Skip to content

Architecture

This page describes the comparative study and its arms. For the current clinical product, see the BiobadamexAI workflow.

This page doesn’t exist in any single repo because none of the three has the full picture — each documents its own service, not how it fits with the other two. This is that map.

The code defines three “harnesses” (STUDY_HARNESSES = ["plain", "rheumaai", "grader"], pokta-care-monorepo/apps/api/src/services/biobadamex-study-harness.ts). A study “arm” is really (model × harness × optional prompt) — there’s no fixed hardcoded model list, it’s configured at runtime via the STUDY_ARMS environment variable.

Harness How it’s invoked Where the code lives
plain LLM call in-process, inside the monorepo — no external repo involved pokta-care-monorepo (plainInvoker)
rheumaai HTTP to RHEUMAI_STRUCTURED_URL external service: rheuma-ai-bioagent, route POST /v1/study/rheumaai-reply
grader HTTP to GRADER_STRUCTURED_URL external service: pokta-grader-bioagent, route POST /v1/chat/completions

The fourth arm (LLM tuned by another rheumatologist) isn’t a separate harness. It’s the plain harness running with a different promptId (rheum-v1), registered in code but with the prompt empty/reserved — if invoked, the code throws instead of silently running with wrong content. It’s not “almost ready,” it’s not built yet.

rheumaai and grader never call each other or share a process — each is an independent HTTP service, deployed separately (Railway), with its own API-key auth. The monorepo is the only piece that knows all three exist; neither service knows about the other (with one cosmetic exception: a comment in rheuma-ai-bioagent’s rheumaai-reply-route.ts explicitly acknowledges its auth technique was copied from pokta-grader-bioagent’s structured.ts — copied, not imported; the repos have no runtime dependency).

Actual sequence from pokta-care-monorepo/apps/api/scripts/biobadamex-sweep-run.ts:

  1. Load the corpus — lists the .docx files in CORPUS_DIR (a filesystem path, never inside a git repo). Derives a note_id = sha256(filename)[:12]; the real filename (which is PHI — the patient’s name) is never logged or persisted.
  2. Resolve the arms — reads STUDY_ARMS from the environment. If empty, exits without touching anything.
  3. Optional dry-run mode — with PLAN_ONLY=1, prints the plan (notes × arms) and exits — zero provider calls, zero DB writes.
  4. Guards against undrained previous runs, then creates a study_run.
  5. For each note: reads/parses it and builds a job per (note, arm) — enqueued in pg-boss.
  6. Workers pick up jobs and call the right invoker for each arm, with a per-provider concurrency semaphore.
  7. Polls every 5s until every cell is “settled” (there’s an output row, or retries are exhausted), up to 45 min.
  8. Final report: counts expected/settled/output/failures, grouped by failure cause.

compareField (biobadamex-eval.ts) compares each field with a four-state verdict, always using an explicit “unknown” check (never a falsy check — so a real 0 is never confused with “missing”):

  • match — both unknown, or the values match (for patient.sex specifically, via an HL7 AdministrativeGender-style catalog, not exact string comparison).
  • mismatch — both known, values differ.
  • missing — the gold has a value, the extraction doesn’t.
  • unexpected — the extraction has a value, the gold doesn’t.

“Completeness” is a presence rate, not an accuracy rate. It’s present / total over the required fields that aren’t UNKNOWN — it doesn’t factor in whether the value is correct against the gold. Accuracy is a separate metric (variant_accuracy, delta_grader_minus_plain). This distinction already caused a real bug in the study (arm-2 structuring the model’s prose instead of the original note inflated completeness without improving accuracy) — if you’re reporting a number, be explicit about which of the two metrics it is.

  • biobadamex_study_arm_outputs — one row per (run_id, note_id, arm_id). Relevant columns: model_id, harness, record (the extracted JSON), das28, missing_required, latency_ms, tokens. No FK on note_id — on purpose.
  • biobadamex_study_analysis — aggregated only, long format, no note-level key (on purpose, to prevent re-identification). Columns: run_id, arm_id, model_id, harness, scope, field_key, metric, numerator, n, value, confidence intervals. This is where you filter by scope="field", metric="field_present_rate" to see the field-by-field breakdown (e.g. the sex field).

Lives outside any repo, at a filesystem path passed via CORPUS_DIR. The census script (biobadamex-instrument-census.ts) explicitly verifies that path is not inside a git working tree — “Real patient notes must never be committable” is a literal comment in the code, not just a convention. See Guardrails for the rules on working with synthetic data instead.