Outcome-Verified Clinical Intelligence
Clinical AI That
Proves Itself
Amphibia monitors patients continuously — in sickness and in health — generates differential diagnoses from that living record, shows its reasoning, and tracks whether it was right, confirmed by clinicians visit after visit.
Every other clinical AI asks for your trust. Amphibia earns it, one verified outcome at a time.
The Accountability Gap
Healthcare Bought AI.
Nobody Can Prove It Works.
Health systems spent billions on scribes and knowledge tools. None of them can answer the only question that matters: was the AI right?
The current generation
Document the visit. Never ask if the assessment was right.
Answers from literature. Never touches the patient's chart.
Superhuman on exam questions. Unmeasured on real patients.
Store and report. No reasoning, no accountability.
The evidence says the same
Independent evaluations of clinical AI find reduced burnout — but no demonstrated impact on patient outcomes.
Diagnostic accuracy is claimed on benchmarks, not measured in production.
No tool closes the loop between AI suggestion and confirmed patient outcome.
Amphibia was built to close that loop: every suggestion is reviewed by a clinician and verified against what actually happened to the patient.
The Amphibia Approach
The Whole Consultation,
Not One Step
Triage, differential, which study to order next, what a result rules out, and a plan the patient can actually follow — reasoned from a record that accumulates between visits, and held accountable for every suggestion.
Reasons Within Real Constraints
Coverage, formulary, out-of-pocket ceiling, literacy, distance to care. A plan the patient cannot afford or obtain is not a plan, and most systems never ask.
Triage the Model Cannot Lower
Deterministic rules set the floor from recorded vitals. The model may escalate; it can never de-escalate below what the measurements establish.
Studies Ranked by What They Settle
Ordered by discriminating value and by yield per unit spent, not by thoroughness — then which candidates each result rules out, and which it only makes less likely.
Treatment That Travels
Every prescribed drug resolved to its active ingredient and ATC code, translated to the pharmacy shelf of the patient's country, and screened against what they already take.
The Interval Nobody Sees
A nightly sweep for the value that has been out of range with no change in therapy. That interval, repeated, is what becomes the expensive complication.
Calibrated, and Honest About It
Confidence is capped at what validation supports. On its target population the system understates its own reliability — the only acceptable direction for the error.
The Evidence Loop
Every Suggestion
Is Held Accountable
Benchmarks measure models on exam questions. The Evidence Loop measures Amphibia on your patients — and shows you the results.
Consultation
Voice or text. The full encounter, transcribed and structured.
Differential, With Reasoning
Ranked diagnoses generated from the live chart — every step of the reasoning recorded.
Clinician Decision
The physician reviews, agrees or overrides. Authority never leaves the human.
Confirmed Outcome
Post-visit, the clinician confirms what was actually right. The loop closes.
Verified Performance
Acceptance rate, real NNT, calibration — measured per clinician, in production.
Assessments
128
Reviewed
94
Acceptance
71%
Real NNT
1.4
Illustrative interface. Metrics are computed per clinician from confirmed post-visit outcomes.
View live evidence registry →Continuous, Not Episodic
Healthcare Shows Up When
You're Sick. Amphibia Doesn't Leave.
Most clinical AI activates for fifteen minutes inside the exam room. Amphibia monitors the patient across the entire health lifecycle — baseline vitals, self-reported journals, treatment response, long-term stability — so every diagnosis reasons from years of context, and every outcome feeds back into the record.
Healthy
Baseline monitoring
At Risk
Early detection
Diagnosed
Coordinated reasoning
Treated
Outcome tracking
Stable
Long-term prevention
One continuous intelligence layer from healthy to stable — not five disconnected tools that only see the patient when something is already wrong.
Built for the Point of Care
For Clinicians.
For Health Systems.
Amphibia serves the people who make diagnostic decisions and the institutions accountable for them.
Clinicians
Differential diagnoses from the live chart, reasoning you can audit, and a personal performance dashboard showing how the AI does with your patients.
Health Systems
Outcome-verified AI performance across your clinician population, full audit trails on every inference, and FHIR-native record integration.
One intelligence layer, every stakeholder
The same continuously collected record flows — with patient consent and a full audit trail — to everyone the health system depends on.
Clinical Trust Architecture
Healthcare AI Requires
Explainability
Built for regulated environments where trust, transparency and human oversight are non-negotiable — and where a declared safeguard that cannot refuse is documentation, not a control.
Decision Support, Not Replacement
Amphibia enhances clinical judgment — it never substitutes for it. Every recommendation requires human review, and a final diagnosis cannot be built on a suggestion nobody reviewed.
Explainable Reasoning
No black boxes. Every insight includes reasoning clinicians can audit, question, and override.
Safeguards That Refuse
Layers that consult no model at all: they decline outside the validated population, decline on a record too thin to reason from, and raise alerts from the measurements whether or not the model noticed.
Complete Auditability
Full decision provenance. Every recommendation, data point, and access — logged and traceable.
Population Scale, Across Borders
Health Intelligence
at Population Scale
Infrastructure for governments and health systems to make evidence-based decisions across entire populations. Built on FHIR, the international interoperability standard, and multilingual from day one — the same intelligence layer works across health systems.
Early Risk Detection
Identify at-risk cohorts before conditions manifest clinically.
Preventive Interventions
Enable proactive programs informed by continuous population intelligence.
Resource Optimization
Allocate healthcare resources based on real-time population health data.
Epidemiological Insights
Real-time surveillance and trend analysis for public health decision-making — interoperable across institutions and borders via FHIR.
A New Category
Beyond Scribes
and Search
The current generation of clinical AI documents, retrieves, or predicts. None of it is accountable for being right.
Trust & Governance
Compliance You Can
Verify, Not Just Read
We claim only what is running in production. Auditability is not a roadmap item at Amphibia — it is how the system is built.
Every access to health data and every AI inference is logged with actor, purpose and justification — implemented today, not promised.
Each suggestion stores its reasoning steps, the inputs considered and the confidence calculation, in a trace that cannot be edited after the fact.
Clinician–patient access is explicit and verified server-side on every call. No access record, no inference.
Declared performance thresholds are computed from live data, and a breach opens an adverse-event record without anyone asking. On its first run it found one.
Code of Ethics
Commitments that are
enforced, not declared
A code of ethics a system can violate without noticing is a press release. Each commitment below is a rule in the code that runs before the model does — deterministic, auditable, and impossible to talk out of.
The clinician decides. Always.
Amphibia suggests; it does not prescribe, close a case, or act on its own. And when a suggestion reaches a patient without any clinician reviewing it, that is no longer a copilot — the platform's own surveillance flags it and says so.
It never claims certainty.
Stated confidence is capped at 0.85 by a rule outside the model, not by the model's own judgement. Overconfidence induces premature closure; a ceiling induces verification.
It stays silent when it does not know.
Below three elements in the record it issues no diagnostic suggestion at all. An opinion built on nothing is worse than no opinion, because it looks like one.
There are uses it refuses.
It does not opine on patients under 18 — outside the validated population — and it rejects any query whose purpose is coverage, eligibility or denial. A clinical system must never become the instrument that withholds care.
It reasons within what the person can actually do.
Coverage, income, schooling, distance to care. A plan the patient cannot afford or reach is not a plan — it is a way of moving the blame onto them.
Every access to health data is recorded.
Without exception, and including the public demo. An audit trail with exceptions is not an audit trail.
Population figures never expose a person.
Aggregates are k-anonymised at the database, and only patients who granted consent are counted. The threshold is enforced by the query itself, not by whoever writes the report.
Our own numbers are measured against us.
Synthetic and demo accounts are excluded from every governance indicator by construction, so no demonstration can ever flatter a real metric. When verification lowers a figure, the lower figure is what gets published.
Every one of these is a line of code you can point at, not a value we hope to hold. Where a commitment is not yet enforced, we say so rather than listing it here.
The 2030 Thesis
Every AI Will Claim
Superhuman Accuracy
By 2030, frontier models will be commodities and regulators will demand performance evidence from production, not exams. Only one kind of clinical AI will matter: the kind that can prove itself on your patients.