Skip to content

Repository files navigation

ResearchOS

ResearchOS is an AI-native workspace for a connected research record. Papers become claims pinned to exact quoted passages; claims become an evidence-bearing graph; verified evidence becomes hypotheses, experiment plans, computed results, and a replayable history. It is built for the OpenAI Build Week Education track.

The five-surface MVP is Literature, Knowledge Graph, Hypotheses, Experiment Planner, and Research Replay. There is deliberately no chat window: AI produces bounded, structured proposals and the researcher verifies the research record. Proposed content is amber, verified evidence is green, and contradictions are red dashed edges.

The included hydrogen-storage pack runs a real ASE + MACE-MP-0 substitution-energy screen in a dedicated CPU worker. Mock runs are visibly watermarked and cannot become hypothesis evidence. This is screening-grade physics, not a substitute for higher-fidelity validation.

Why I built this

I come from a chemical engineering background, and for a long time my research looked the same every time: a folder of PDFs, a document full of half-remembered findings, and no real way to see how any of it fit together. When two papers disagreed on a number, I usually noticed weeks later if at all. And once I reached a conclusion, I couldn't retrace the steps that got me there.

Three things I kept wishing for became the three things ResearchOS is built around.

Everything in one tree. In the lab you carry a hundred facts about a material in your head its enthalpy, its capacity, which catalyst shifts which property and they never live in one place. The knowledge graph puts every claim, material, and property on a single canvas, so a relationship, or a contradiction, becomes something you see rather than something you hope to remember. That red dashed edge where a calculated value disagrees with an experiment is the exact thing I used to miss.

A replay of my own reasoning. Research isn't just the answer; it's the path. Because every action here is an event, I can scrub the whole project back to the first question and watch the record rebuild itself every claim, every verification, every result, with the source behind each one. Nothing gets lost, and I can hand the trail to someone else.

Assistance I can actually trust. I wanted help reading and connecting papers without giving up judgment. So the AI only ever proposes shown in amber and nothing counts until I check it against the exact quote and it turns green. The human stays in the loop, every claim carries accountability, and the model's speed serves my scrutiny instead of replacing it.

ResearchOS is the tool I wanted while doing the work: one place where the evidence, the reasoning, and the history stay together and show their work.

Architecture

Next.js web (five folded surfaces)
        │ HTTP + SSE
        ▼
FastAPI ───────────────► Postgres 16 + pgvector
   │                         │ events outbox / provenance
   ├── Redis + arq ─► default worker (AI jobs)
   └── Redis + arq ─► compute worker (ASE + MACE CPU)
                              │
                    hydrogen-storage domain pack

Setup

Requirements: Docker Compose v2. Python and Node are only needed for local development outside containers.

cp .env.example .env
# Add OPENAI_API_KEY to .env when using extraction, hypothesis, planning, or interpretation.
docker compose up --build

Open http://localhost:3000. API health is at http://localhost:8000/health. PostgreSQL publishes on host port 5433; container-to-container traffic remains postgres:5432.

Seed the demo project after the stack is running:

make seed

The seed command ingests every PDF in seed/papers/ and accepts optional arXiv IDs from seed/papers/papers.txt. It is idempotent for already-ingested file titles.

Demo walkthrough

  1. Start in Literature. Open one of the seeded papers, run extraction with a configured key, and verify claims with V; reject with R; edit with E; navigate with J/K.
  2. Open Knowledge Graph. Proposed edges are amber; verification turns them green. A numeric disagreement appears as a red dashed contradiction edge. Click it to inspect both source-backed claims.
  3. In Hypotheses, draft up to three candidates from the selected graph region. Each must show evidence for and against, or clearly warn that the project may be one-sided.
  4. Use Experiment Planner to create a substitution-energy screen. Run MACE for real SSE logs, then inspect the computed bars, verified-literature bands, and compatible hypothesis threshold. Mock output remains watermarked.
  5. In Research Replay, scrub to zero and use Papers, Graph, Hypothesis, and Results to rebuild the same project history. Click an event, graph object, or conflict to inspect its provenance.

GPT-5.6 usage and grounding

TIER_EXTRACTION defaults to gpt-5.6-terra for parallel structured claim extraction. TIER_REASONING defaults to gpt-5.6-sol for hypothesis drafting, constrained plan filling, and result interpretation. TIER_FAST defaults to gpt-5.6-luna for entailment verification and relevance triage. EMBED_MODEL defaults to text-embedding-3-large for chunk embeddings.

Grounding is enforced server-side: quoted spans must occur in the referenced chunk after whitespace-only normalization, with mapped original-text offsets retained for highlighting; numeric values are re-parsed from the quote; hypothesis and interpretation references are checked against project records; and unverified claims cannot become hypothesis evidence. Every model call is persisted as a trace with prompt version, model, token counts, latency, output, and validation outcome.

How Codex built this

Codex implemented the scaffold, schema, ingestion pipeline, extraction/verification plumbing, graph materializer, planner, MACE worker integration, result comparison, and replay projection. I directed product direction, scientific constraints, prompt intent, verification interaction design, physics validation, and demo choreography.

Canonical Codex /feedback session ID: 019f5ff1-8c19-73c1-aaf2-1a5e7ff5b8e8.

FAQ

Isn't this RAG with extra steps? RAG answers and forgets. ResearchOS builds a persistent, human-verified, replayable knowledge structure and closes the loop with real computed physics.

How do you prevent hallucination? Enforced, not requested: every claim must quote a verbatim source span, numeric values are re-parsed deterministically, entailment is independently checked, references are validated, and unverified content is amber and excluded from hypothesis evidence.

Won't students cheat with this? Every object carries provenance and verification state. AI content remains visibly marked until a human owns it, giving instructors more inspectability rather than less.

Why hydrogen storage? It is an urgent clean-energy research area with a tight literature and physics cheap enough to run live. The core schema is domain-agnostic; hydrogen storage is domain pack one.

What is real versus mocked? Real: PDF/arXiv ingestion, GPT-5.6 structured-output jobs when configured, MACE-MP simulation, provenance, and replay. Mock adapter output is visibly labeled. Not built: cluster adapter, multi-user collaboration, and additional workflow templates.

Limitations

Figure and table-image extraction are deliberately excluded because they are a high-hallucination source. MACE ML-potential results are screening-grade and must not be treated as experimental truth. The graph knows only imported material; this MVP is single-user, ships one workflow template, and displays proposed claims in amber before verification, while retaining the verified-only evidence gate for hypotheses.

Development checks

docker compose run --rm --no-deps api pytest -q
cd apps/web && npm run typecheck && npm run build

License

MIT. See LICENSE.

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages