How it works

One picture of what medicine knows, and how it is built.

OpenQuestion reads the public record of medical research, keeps one graph of what it says, attaches a probability to every claim, and points at the experiment that would change that probability most. This page describes each step and where it is still rough. If you want to learn to read the charts first, start with the tutorial.

One graph, pinned releases

Everything on the site is a view of a single graph: entities such as drugs, diseases, outcomes, trials and papers, and the relationships between them. The graph is rebuilt in generations, and each generation advances in numbered epochs as new source records are folded in. Every page names the generation and epoch it was read from, so a number you saw last month can be traced back to the exact release that produced it.

The current release covers cardiometabolic medicine and diabetes: about 4,000 registered trials and the papers linked to them. Other specialties are built the same way and will appear as they are published.

Sources and rights

Records come from public catalogs: PubMed and PubMed Central for the literature, ClinicalTrials.gov for registered trials and their posted results, OpenAlex for bibliography and citations, NIH RePORTER for funding, and around sixty biological and clinical databases for the entities themselves. The full list, with record counts, is on the Sources page, and live loading state is on Coverage.

Loading a source is not the same as using it. Each source passes a separate rights review before its records can be shown or can contribute relationships to a published release. Where a source allows metadata but not full text, only the metadata is used, and the page says so.

The evidence map

The map is a table. Each row is an intervention class, such as SGLT2 inhibitors or statins. Each column is an outcome domain, such as glycemic control or cardiovascular events. Every registered trial in the release is placed in the cells for the interventions it tested and the outcomes it measured, using rules that map trial labels to classes and domains. The current release places about 95 percent of trials this way; the rest carry labels the rules do not yet recognise and are listed but not placed.

Risk-factor rows sit above the interventions that act on them, so a column reads as a causal chain: what the population evidence says about the exposure, then what trials say about treating it. A blank cell is a gap, and blank cells are as much the point of the map as full ones.

Reading results

A placed trial says what was tested. Whether it worked comes from two places. Posted results on ClinicalTrials.gov give per-analysis estimates with their intervals and p values. Published abstracts give effect estimates, which are extracted when the sentence names the drug and compares it to something, and are ignored when they describe an association or a within-group change.

Each estimate is oriented so that the intervention is always compared against the comparator, no matter the order the sponsor entered the groups, and is given a polarity so that lower is better for harms and higher is better for benefits. Only then can two trials be compared. An estimate that cannot be oriented or given a polarity is shown as unreadable rather than guessed.

Votes and evidence families

For every claim in a cell, for example that a drug class lowers cardiovascular events against placebo, each trial casts one vote: supports, contradicts, or inconclusive. Primary analyses decide the vote. Secondary analyses can support but never contradict on their own. A trial with no readable result does not vote.

Reports of the same study are grouped into one evidence family before voting, so a registry entry, its primary publication and three secondary papers count once. Publications with no registry link form their own family. This is what stops a well-published trial from outvoting a better one.

Belief

Votes become a belief: a probability that the claim is true, with a band around it and a state. One supporting family gives a belief near one half, marked unreplicated. Independent agreeing families move it up and narrow the band, marked replicated. Disagreement widens the band and marks the claim contested, unless the dissent is a small minority against several supporting families, in which case it lowers the belief without changing the state.

Each family's contribution is recorded, and every paper page shows the belief with and without that paper, so you can see exactly how much any one study is carrying.

Gaps and the next experiment

A gap is something the graph knows it does not know: a claim with a single family, a contested claim, a cell with trials but no readable result, or a cell with none at all. For each gap the graph proposes an experiment and estimates its expected information gain, on a scale where one would settle the question, from how far the belief would move if the result confirmed or contradicted the claim.

Proposals are ranked by that gain, weighted by the stakes of the outcome and the prevalence of the condition, divided by the likely cost of the trial, with extra weight when no company has a commercial reason to run it and less when a trial is already recruiting. The ranked list is the test-next view of the map.

Compiled questions

A question page is a compiled slice of the graph: the entities in scope, the claim, its belief and lineage, the source records behind it, the open gaps and the proposed experiment. Each is pinned to one generation and epoch and carries a version history. The box on the home page matches what you type against these compiled questions by shared terms. It never generates an answer.

Before a question is compiled its claims are checked against the source records. Human risk-of-bias review of the included studies is planned but has not started, and every question says so until it has.

Known limits

  • Population. Cell-level claims are not scoped by population unless a condition filter is applied. A metformin trial in polycystic ovary syndrome counts toward the same cell as one in type 2 diabetes.
  • Trial size. Families are unweighted. A trial of three hundred people casts the same vote as one of thirty thousand.
  • Non-inferiority. The registry does not record a trial's hypothesis. A head-to-head trial that succeeded by showing non-inferiority can currently read as a vote against.
  • Full text. Most published estimates are read from abstracts. Only a few hundred papers in the release have machine-readable results sections, and tables are not yet parsed.
  • Background therapy. When a drug is given to both arms, comparisons involving it are unclear. About half of metformin comparisons fall here.

What we commit to

Every number links to the records that produced it. Every page names its release. Nothing on the site is written by a language model. Beliefs are stated before new results arrive and are scored when they land, and the scoring will be published on the site. When a method changes, the release number changes with it.

Questions and corrections: hello@socratic.science.