The AI-native EHR, graded by completeness.
We're building the electronic health record defined by what it captures, not what it looks like. Rheumatology is the proving ground: a multi-arm, multi-model benchmark of PoktaGrader against real, physician-certified clinical notes.
Prefer email? Write directly to info@poktalabs.com.
- 110
- physician-certified notes benchmarked
- 220
- extractions run, 0 failures
- <$0.20
- total compute for the sweep
- 93.6% / 94.5%
- diagnostic accuracy, baseline vs. PoktaGrader
Healthcare has more guidelines than time to apply them.
Regulation is forcing the record to exist
Mexican health authorities have begun enforcing electronic health record requirements, in a market where many specialists still keep patient histories on paper or in a word processor. That is regulator-created demand, not a feature we have to sell.
Guidelines are public. Attention is not.
Clinical protocols for most conditions are published, peer-reviewed, and free. What is scarce is the fifteen minutes a specialist has to apply all of them, mid-consultation, while also treating the patient in front of them.
Models can finally read a chart
Structured extraction and clinical reasoning over messy, real-world notes, inconsistent formats, incomplete fields, shorthand that assumes a career of context, are exactly the failure mode general-purpose language models have closed over the last two years.
Latin America is the greenfield
Legacy EHR vendors were built for US billing codes and insurance workflows. They have little reason to localize for Mexican specialty medicine. A market this underserved, with a regulatory tailwind this strong, does not come along often.
Defined by what's missing, not just what's stored.
Most electronic records are a database with a form on top: they store whatever a clinician types, complete or not. PoktaGrader inverts that. Every note is measured against the completeness a disease actually requires, so gaps surface at the moment they can still be closed, not the moment an audit finds them.
The moat isn't the checklist, clinical guidelines are already public. The moat is knowing which missing item is worth interrupting a consultation for, and when.
A factorial study, not a demo.
We test PoktaGrader the way we'd want a competitor's claims tested: against a plain extraction baseline, a general clinical-reasoning assistant, and a prompt-only control, all run across the same real notes and, increasingly, the same slate of models.
- 01
Baseline extraction
A plain model call, no tools, no persona. The floor every other arm has to beat.
- 02
RheumaAI
Pokta Labs' own general clinical-reasoning assistant for rheumatology, run with its own tools and retrieval.
- 03
PoktaGrader
The completeness engine itself: its own prompt, its own tools, its own view of what a complete note requires.
- 04
Prompt-only control
A third-party specialist's own prompt, no retrieval. Isolates how much of any result is just a better prompt.
Every note is real: drawn from an active rheumatology practice, already treated by its physician as clinically complete. Results are reported at the corpus level only, never per clinician, and this research carries no official endorsement or affiliation from any medical college, registry, or institution.
Honest numbers, including the ones that don't flatter us.
110
physician-certified notes benchmarked
64 rheumatoid arthritis · 46 spondyloarthritis
220
extractions run, 0 failures
full corpus × 2 methods, complete provenance
<$0.20
total compute for the sweep
the whole corpus, not a sample
93.6% / 94.5%
diagnostic accuracy, baseline vs. PoktaGrader
both correctly read the disease from the note
Why a tie is the finding
This isn't a highlight reel. We pre-registered our primary endpoint before running the numbers, specifically so we couldn't pick the metric that flatters our own arm after the fact. On the first full sweep, PoktaGrader trailed the baseline by 4.2 completeness points, traced to two fields still being tuned. The corrected sweep closed that gap to a completeness delta statistically indistinguishable from zero, 95% CI: −1.4 to +1.1 points, at matched 93.6% / 94.5% diagnostic accuracy. A benchmark that catches and diagnoses its own miss is the kind of result that publishes.
Preliminary results from ongoing research, shared for transparency. Figures describe completeness, whether data was captured, not yet clinical accuracy on every field. Reporting follows two pre-declared analyses: worst-case imputation as the headline, complete-case intersection as the arm-vs-arm secondary.
Model-agnostic by design, running on Nebius.
- gpt-oss-120b
- gpt-5.6-sol
- Kimi K3
- GLM-5.1
- MiniMax M3
- Claude
Every arm above runs on gpt-oss-120b, an open-weight model, as the hard baseline for the whole study, chosen so the result doesn't depend on any one vendor's model card. We're extending the same benchmark across a wider slate, run through the identical harness with cross-model fallback disabled, so every cell in the comparison is exactly what it claims to be.
That only works at this pace because of Nebius. Nebius' AI cloud is our inference layer for the open and frontier models in the study, giving us the throughput to run a real multi-model factorial design instead of a single-vendor demo, at a cost low enough that the entire 110-note, 220-extraction sweep so far has cost under $0.20 in compute.
From a benchmark to an AI-native EHR.
Reference labels and more models
A blind, field-level accuracy subset, and the same benchmark run across the full model slate.
Prospective interruption-quality logging
Capturing which missing-field prompts a clinician actually acts on, in real time. This is the real moat.
AI-native EHR
A record defined by completeness from the first note, not retrofitted onto one.
Already built
Built
- Corpus fixed
- Disease-specific instruments defined
- Methods built
Run
- Baseline model fixed
- Full corpus swept
- Grader evaluated and corrected
Three ways to work with us.
Sponsors & research funders
Fund the next phase: the full multi-model sweep, and the prospective interruption-quality study that turns a retrospective benchmark into a live clinical tool.
Investors
An AI-native EHR platform, in a market with a regulatory tailwind and legacy vendors that haven't localized, grounded in a published, pre-registered benchmark instead of a pitch deck.
Collaborators & contributors
Clinicians to help define completeness instruments beyond rheumatology. Researchers and engineers to extend the harness, the model slate, and the open tooling.
Who's behind this.
Ing. Angel Melendez C.
Founder, Pokta Labs
Dr. Erick Zamora T.
Rheumatologist, clinical lead
Let's build the record healthcare actually needs.
Whether you're funding research, evaluating an investment, or want to bring your specialty into the benchmark, we'd like to talk.
A research initiative from Pokta Labs, or write directly to info@poktalabs.com.
- 93.6% / 94.5%
- diagnostic accuracy, baseline vs. PoktaGrader
- <$0.20
- total compute for the full sweep, no sample