Architecture
This page describes the comparative study and its arms. For the current clinical product, see the BiobadamexAI workflow.
This page doesn’t exist in any single repo because none of the three has the full picture — each documents its own service, not how it fits with the other two. This is that map.
The three arms, who runs them
Section titled “The three arms, who runs them”The code defines three “harnesses” (STUDY_HARNESSES = ["plain", "rheumaai", "grader"], pokta-care-monorepo/apps/api/src/services/biobadamex-study-harness.ts). A study “arm” is really (model × harness × optional prompt) — there’s no fixed hardcoded model list, it’s configured at runtime via the STUDY_ARMS environment variable.
| Harness | How it’s invoked | Where the code lives |
|---|---|---|
plain |
LLM call in-process, inside the monorepo — no external repo involved | pokta-care-monorepo (plainInvoker) |
rheumaai |
HTTP to RHEUMAI_STRUCTURED_URL |
external service: rheuma-ai-bioagent, route POST /v1/study/rheumaai-reply |
grader |
HTTP to GRADER_STRUCTURED_URL |
external service: pokta-grader-bioagent, route POST /v1/chat/completions |
The fourth arm (LLM tuned by another rheumatologist) isn’t a separate harness. It’s the plain harness running with a different promptId (rheum-v1), registered in code but with the prompt empty/reserved — if invoked, the code throws instead of silently running with wrong content. It’s not “almost ready,” it’s not built yet.
Request flow
Section titled “Request flow”flowchart TD
M["pokta-care-monorepo<br/>(study orchestrator)"]
P["plain<br/>(in-process LLM call)"]
R["rheumaai<br/>HTTP POST → rheuma-ai-bioagent<br/>/v1/study/rheumaai-reply"]
G["grader<br/>HTTP POST → pokta-grader-bioagent<br/>/v1/chat/completions"]
O[("biobadamex_study_arm_outputs<br/>one row per note × arm")]
S["scoring<br/>(biobadamex-eval.ts)"]
A[("biobadamex_study_analysis<br/>aggregated, no per-note key")]
M --> P
M --> R
M --> G
P --> O
R --> O
G --> O
O --> S
S --> Arheumaai and grader never call each other or share a process — each is an independent HTTP service, deployed separately (Railway), with its own API-key auth. The monorepo is the only piece that knows all three exist; neither service knows about the other (with one cosmetic exception: a comment in rheuma-ai-bioagent’s rheumaai-reply-route.ts explicitly acknowledges its auth technique was copied from pokta-grader-bioagent’s structured.ts — copied, not imported; the repos have no runtime dependency).
Sweep flow — start to finish
Section titled “Sweep flow — start to finish”Actual sequence from pokta-care-monorepo/apps/api/scripts/biobadamex-sweep-run.ts:
- Load the corpus — lists the
.docxfiles inCORPUS_DIR(a filesystem path, never inside a git repo). Derives anote_id = sha256(filename)[:12]; the real filename (which is PHI — the patient’s name) is never logged or persisted. - Resolve the arms — reads
STUDY_ARMSfrom the environment. If empty, exits without touching anything. - Optional dry-run mode — with
PLAN_ONLY=1, prints the plan (notes × arms) and exits — zero provider calls, zero DB writes. - Guards against undrained previous runs, then creates a
study_run. - For each note: reads/parses it and builds a job per
(note, arm)— enqueued inpg-boss. - Workers pick up jobs and call the right invoker for each arm, with a per-provider concurrency semaphore.
- Polls every 5s until every cell is “settled” (there’s an output row, or retries are exhausted), up to 45 min.
- Final report: counts expected/settled/output/failures, grouped by failure cause.
Scoring — what “completeness” means
Section titled “Scoring — what “completeness” means”compareField (biobadamex-eval.ts) compares each field with a four-state verdict, always using an explicit “unknown” check (never a falsy check — so a real 0 is never confused with “missing”):
- match — both unknown, or the values match (for
patient.sexspecifically, via an HL7 AdministrativeGender-style catalog, not exact string comparison). - mismatch — both known, values differ.
- missing — the gold has a value, the extraction doesn’t.
- unexpected — the extraction has a value, the gold doesn’t.
“Completeness” is a presence rate, not an accuracy rate. It’s present / total over the required fields that aren’t UNKNOWN — it doesn’t factor in whether the value is correct against the gold. Accuracy is a separate metric (variant_accuracy, delta_grader_minus_plain). This distinction already caused a real bug in the study (arm-2 structuring the model’s prose instead of the original note inflated completeness without improving accuracy) — if you’re reporting a number, be explicit about which of the two metrics it is.
Where results land
Section titled “Where results land”biobadamex_study_arm_outputs— one row per(run_id, note_id, arm_id). Relevant columns:model_id,harness,record(the extracted JSON),das28,missing_required,latency_ms, tokens. No FK onnote_id— on purpose.biobadamex_study_analysis— aggregated only, long format, no note-level key (on purpose, to prevent re-identification). Columns:run_id,arm_id,model_id,harness,scope,field_key,metric,numerator,n,value, confidence intervals. This is where you filter byscope="field",metric="field_present_rate"to see the field-by-field breakdown (e.g. thesexfield).
Real note corpus
Section titled “Real note corpus”Lives outside any repo, at a filesystem path passed via CORPUS_DIR. The census script (biobadamex-instrument-census.ts) explicitly verifies that path is not inside a git working tree — “Real patient notes must never be committable” is a literal comment in the code, not just a convention. See Guardrails for the rules on working with synthetic data instead.