Outcome-Verified Clinical Intelligence

Clinical AI That
Proves Itself

Amphibia monitors patients continuously — in sickness and in health — generates differential diagnoses from that living record, shows its reasoning, and tracks whether it was right, confirmed by clinicians visit after visit.

Every other clinical AI asks for your trust. Amphibia earns it, one verified outcome at a time.

The Accountability Gap

Healthcare Bought AI.
Nobody Can Prove It Works.

Health systems spent billions on scribes and knowledge tools. None of them can answer the only question that matters: was the AI right?

The current generation

01
Ambient scribes

Document the visit. Never ask if the assessment was right.

02
Medical knowledge AI

Answers from literature. Never touches the patient's chart.

03
Benchmark models

Superhuman on exam questions. Unmeasured on real patients.

04
EHR analytics

Store and report. No reasoning, no accountability.

The evidence says the same

Independent evaluations of clinical AI find reduced burnout — but no demonstrated impact on patient outcomes.

Diagnostic accuracy is claimed on benchmarks, not measured in production.

No tool closes the loop between AI suggestion and confirmed patient outcome.

Amphibia was built to close that loop: every suggestion is reviewed by a clinician and verified against what actually happened to the patient.

The Amphibia Approach

The Whole Consultation,
Not One Step

Triage, differential, which study to order next, what a result rules out, and a plan the patient can actually follow — reasoned from a record that accumulates between visits, and held accountable for every suggestion.

01

Reasons Within Real Constraints

Coverage, formulary, out-of-pocket ceiling, literacy, distance to care. A plan the patient cannot afford or obtain is not a plan, and most systems never ask.

02

Triage the Model Cannot Lower

Deterministic rules set the floor from recorded vitals. The model may escalate; it can never de-escalate below what the measurements establish.

03

Studies Ranked by What They Settle

Ordered by discriminating value and by yield per unit spent, not by thoroughness — then which candidates each result rules out, and which it only makes less likely.

04

Treatment That Travels

Every prescribed drug resolved to its active ingredient and ATC code, translated to the pharmacy shelf of the patient's country, and screened against what they already take.

05

The Interval Nobody Sees

A nightly sweep for the value that has been out of range with no change in therapy. That interval, repeated, is what becomes the expensive complication.

06

Calibrated, and Honest About It

Confidence is capped at what validation supports. On its target population the system understates its own reliability — the only acceptable direction for the error.

The Evidence Loop

Every Suggestion
Is Held Accountable

Benchmarks measure models on exam questions. The Evidence Loop measures Amphibia on your patients — and shows you the results.

01

Consultation

Voice or text. The full encounter, transcribed and structured.

02

Differential, With Reasoning

Ranked diagnoses generated from the live chart — every step of the reasoning recorded.

03

Clinician Decision

The physician reviews, agrees or overrides. Authority never leaves the human.

04

Confirmed Outcome

Post-visit, the clinician confirms what was actually right. The loop closes.

05

Verified Performance

Acceptance rate, real NNT, calibration — measured per clinician, in production.

AI Performance — Dr. S. AlvarezOutcome-Verified

Assessments

128

Reviewed

94

Acceptance

71%

Real NNT

1.4

Illustrative interface. Metrics are computed per clinician from confirmed post-visit outcomes.

View live evidence registry →

Continuous, Not Episodic

Healthcare Shows Up When
You're Sick. Amphibia Doesn't Leave.

Most clinical AI activates for fifteen minutes inside the exam room. Amphibia monitors the patient across the entire health lifecycle — baseline vitals, self-reported journals, treatment response, long-term stability — so every diagnosis reasons from years of context, and every outcome feeds back into the record.

Healthy

Baseline monitoring

At Risk

Early detection

Diagnosed

Coordinated reasoning

Treated

Outcome tracking

Stable

Long-term prevention

One continuous intelligence layer from healthy to stable — not five disconnected tools that only see the patient when something is already wrong.

Built for the Point of Care

For Clinicians.
For Health Systems.

Amphibia serves the people who make diagnostic decisions and the institutions accountable for them.

01

Clinicians

Differential diagnoses from the live chart, reasoning you can audit, and a personal performance dashboard showing how the AI does with your patients.

02

Health Systems

Outcome-verified AI performance across your clinician population, full audit trails on every inference, and FHIR-native record integration.

One intelligence layer, every stakeholder

The same continuously collected record flows — with patient consent and a full audit trail — to everyone the health system depends on.

PatientsContinuous monitoring and coordinated care between visits.
Insurers & PayersOutcome-verified, value-based care analytics.
EmployersPrivacy-first workforce wellbeing.
Public HealthPopulation-scale surveillance and insights.
IndividualsProactive monitoring before anything goes wrong.

Clinical Trust Architecture

Healthcare AI Requires
Explainability

Built for regulated environments where trust, transparency and human oversight are non-negotiable — and where a declared safeguard that cannot refuse is documentation, not a control.

01

Decision Support, Not Replacement

Amphibia enhances clinical judgment — it never substitutes for it. Every recommendation requires human review, and a final diagnosis cannot be built on a suggestion nobody reviewed.

02

Explainable Reasoning

No black boxes. Every insight includes reasoning clinicians can audit, question, and override.

03

Safeguards That Refuse

Layers that consult no model at all: they decline outside the validated population, decline on a record too thin to reason from, and raise alerts from the measurements whether or not the model noticed.

04

Complete Auditability

Full decision provenance. Every recommendation, data point, and access — logged and traceable.

Population Scale, Across Borders

Health Intelligence
at Population Scale

Infrastructure for governments and health systems to make evidence-based decisions across entire populations. Built on FHIR, the international interoperability standard, and multilingual from day one — the same intelligence layer works across health systems.

Early Risk Detection

Identify at-risk cohorts before conditions manifest clinically.

Preventive Interventions

Enable proactive programs informed by continuous population intelligence.

Resource Optimization

Allocate healthcare resources based on real-time population health data.

Epidemiological Insights

Real-time surveillance and trend analysis for public health decision-making — interoperable across institutions and borders via FHIR.

A New Category

Beyond Scribes
and Search

The current generation of clinical AI documents, retrieves, or predicts. None of it is accountable for being right.

AI ScribesDocument the encounterNo diagnosis, no outcome tracking
Medical Knowledge AIAnswer from literatureNever touches the patient's chart
Electronic Health RecordsStore and reportNo reasoning, no accountability
AmphibiaDiagnoses from live patient data, shows its reasoning, and verifies itself against confirmed outcomes.

Trust & Governance

Compliance You Can
Verify, Not Just Read

We claim only what is running in production. Auditability is not a roadmap item at Amphibia — it is how the system is built.

HIPAA-Compliant Audit Trail

Every access to health data and every AI inference is logged with actor, purpose and justification — implemented today, not promised.

Reasoning Provenance

Each suggestion stores its reasoning steps, the inputs considered and the confidence calculation, in a trace that cannot be edited after the fact.

Role-Based Access Control

Clinician–patient access is explicit and verified server-side on every call. No access record, no inference.

Surveillance That Runs Itself

Declared performance thresholds are computed from live data, and a breach opens an adverse-event record without anyone asking. On its first run it found one.

Code of Ethics

Commitments that are
enforced, not declared

A code of ethics a system can violate without noticing is a press release. Each commitment below is a rule in the code that runs before the model does — deterministic, auditable, and impossible to talk out of.

01

The clinician decides. Always.

Amphibia suggests; it does not prescribe, close a case, or act on its own. And when a suggestion reaches a patient without any clinician reviewing it, that is no longer a copilot — the platform's own surveillance flags it and says so.

02

It never claims certainty.

Stated confidence is capped at 0.85 by a rule outside the model, not by the model's own judgement. Overconfidence induces premature closure; a ceiling induces verification.

03

It stays silent when it does not know.

Below three elements in the record it issues no diagnostic suggestion at all. An opinion built on nothing is worse than no opinion, because it looks like one.

04

There are uses it refuses.

It does not opine on patients under 18 — outside the validated population — and it rejects any query whose purpose is coverage, eligibility or denial. A clinical system must never become the instrument that withholds care.

05

It reasons within what the person can actually do.

Coverage, income, schooling, distance to care. A plan the patient cannot afford or reach is not a plan — it is a way of moving the blame onto them.

06

Every access to health data is recorded.

Without exception, and including the public demo. An audit trail with exceptions is not an audit trail.

07

Population figures never expose a person.

Aggregates are k-anonymised at the database, and only patients who granted consent are counted. The threshold is enforced by the query itself, not by whoever writes the report.

08

Our own numbers are measured against us.

Synthetic and demo accounts are excluded from every governance indicator by construction, so no demonstration can ever flatter a real metric. When verification lowers a figure, the lower figure is what gets published.

Every one of these is a line of code you can point at, not a value we hope to hold. Where a commitment is not yet enforced, we say so rather than listing it here.

The 2030 Thesis

Every AI Will Claim
Superhuman Accuracy

Benchmark Claims
Verified Outcomes

By 2030, frontier models will be commodities and regulators will demand performance evidence from production, not exams. Only one kind of clinical AI will matter: the kind that can prove itself on your patients.

Don't take a clinical AI's word for it.

Ask it to prove itself.