scinr
AI-powered knowledge extraction for life sciences
scinr (scinr.newton) is a Python library for turning life sciences documents and tabular data into structured, queryable knowledge.
For unstructured documents, scinr extracts domain entities and relationships and stores them as connected knowledge graphs in Neo4j, with optional document storage in MongoDB.
It uses LLMs and Pydantic models to extract structured, domain-specific information from scientific content.
Key Features
-
5-stage document pipeline — Preprocess, extract structure, ingest, annotate, and extract domain entities from unstructured documents.
-
Tabular data pipeline — Normalize scientific spreadsheets and extract structured entities from tabular data.
-
Multi-format ingestion — Supports
.pdf,.docx,.xlsx,.xls, and.csv. Support for.pptx,.json,.xml,.html, and.txtis planned in the roadmap. -
Pydantic extraction models — Define structured schemas for domain entities such as compounds, clinical trials, and assays.
-
Neo4j knowledge graphs — Store extracted entities and relationships with document provenance.
-
Multi-tenant provenance — Stamp every ingested
:Documentwith atenant_id,created_by_user_id, andjob_id, and delete a whole ingestion run in one call withdelete_document(job_id=...). -
Optional MongoDB storage — Store raw files, converted documents, and binary assets using MongoDB/GridFS. Support for other database are planned in the roadmap.
-
LLM-ready documentation — Provides
llms.txtandllms-full.txtfiles for AI coding agents and other LLM-based tools.
Pipelines
Unstructured Documents
Unstructured documents go through five stages:
Raw Documents
(.pdf, .docx, .xlsx, .xls, .csv)
│
▼
1. Preprocess
Convert files to a common representation
│
▼
2. Extraction
Parse sections, structure, and hierarchy
│
▼
3. Ingestion
Store document structure and provenance in Neo4j
│
▼
4. Annotation
Identify relevant content and prepare it for extraction
│
▼
5. Entity Extraction
Extract domain entities and relationships
│
▼
Knowledge Graph
Tabular Data
Tabular data follows a separate pipeline designed for scientific spreadsheets and other structured data:
Tabular Data
(.xlsx, .csv, ...)
│
▼
LLM Entity Extraction
│
▼
Normalization
│
▼
Structured Knowledge
Quick Example
import asyncio
from scinr.newton import configure, run_pipeline
from langchain_aws import ChatBedrockConverse # Or any langchain adapter.
async def main():
llm = ChatBedrockConverse(...)
configure(
llm=llm,
neo4j_uri="bolt://localhost:7687",
neo4j_user="neo4j",
neo4j_password="password",
neo4j_database="neo4j",
mistral_api_key="", # needed for PDF OCR
)
result = await run_pipeline(input_raw="./raw_documents")
print(f"Success: {result.success}")
print(f"Duration: {result.total_duration_seconds:.2f}s")
asyncio.run(main())
How It Works
-
Configure — Set up Neo4j, MongoDB, and LLM providers using
configure()or environment variables. -
Run the pipeline — Call
run_pipeline()with your input directory.scinrhandles document conversion, structure extraction, ingestion, annotation, and entity extraction. -
Query the knowledge graph — Explore extracted entities and relationships in Neo4j.
Documentation
-
Getting Started — Installation, prerequisites, and your first ingestion run.
-
Configuration —
configure(), environment variables, LLM settings, and prompts. -
Architecture — Pipeline stages, data flow, and system design.