5 Things to Know About the New FHIR Server Performance Benchmark

A new public benchmark from Health Samurai, released on June 29, runs HAPI FHIR, Medplum, the Microsoft FHIR Server, and Aidbox on the same bare-metal hardware with the same Synthea dataset. The harness is open source and the dashboard reruns daily. Below are five things worth knowing before any of the numbers get cited in a procurement deck.

1. Four Servers, One Box, Same Rules

The four servers under test are HAPI FHIR, Medplum, the Microsoft FHIR Server, and Aidbox. Each one runs on 8 vCPU and 24 GB of RAM, carved from a single 64-core machine with 500 GB of RAM. Medplum runs as 8 single-vCPU replicas to fit its native scaling model. PostgreSQL 18 backs Aidbox, HAPI, and Medplum; SQL Server 2022 Developer Edition backs the Microsoft FHIR Server.

For broader context, see the FHIR comparison index.

2. CRUD Throughput Has a Wide Spread

The CRUD numbers from the June 29 snapshot are not subtle. Aidbox lands around 5,212 RPS, HAPI around 3,058, Medplum around 1,420, and Microsoft around 440. That is an 11x gap between top and bottom, which matters for any team sizing infrastructure against a real production workload. Honest framing: this is a vendor-run benchmark, since Health Samurai builds Aidbox, but the open repo and daily rerun put any number under review the next day.

The FHIR server vs API gateway comparison covers what raw FHIR throughput actually translates into at the integration layer.

3. Search Performance Is the Other Headline

Search throughput sits closer to the day-to-day shape of EHR traffic than CRUD does. The benchmark reports Aidbox around 3,404 RPS, Medplum around 1,796, HAPI around 1,005, and Microsoft around 261. The notes also call out that Medplum does not support composite search and that Microsoft is slow on quantity and composite parameters. For integration teams that lean on search-heavy workflows, those caveats matter more than the headline number.

The live dashboard with the current snapshot is at the live benchmark dashboard for anyone tracking the daily reruns.

4. Storage Footprints Diverge on the Same Data

After loading the same Synthea dataset of 1,000 patients and roughly 2 million resources, the storage usage tells a story about indexing strategy. Microsoft uses 4.24 GB, Aidbox 6.83 GB, Medplum 11.8 GB, and HAPI 22.6 GB. The 5x spread is not noise. The benchmark notes that HAPI, Medplum, and Microsoft pre-build search indexes on write, while Aidbox ships without default indexes, leaving the operator to choose which ones to add. The trade-off is real either way; it just shows up in different places.

5. The Caveat Section Is Worth Reading

The benchmark is upfront about its limits. The dataset is 1,000 Synthea patients, around 2 million resources, which fits comfortably in memory on a 500 GB host. The empty-database baseline is also acknowledged. The team has signaled that the next post in the series tests at scale, which is the test that production-shaped workloads care about most. For integration architects, the in-memory caveat is the one to keep in mind when projecting the current numbers onto a real production footprint.

The other piece worth reading is the methodology section in the repo. The harness pins each container to 8 vCPU and 24 GB of RAM, runs Synthea ingest in 20 parallel threads, and pushes CRUD load through 300 concurrent k6 threads. Those choices decide what the benchmark can and cannot tell a reader, which is the part most performance reports skip.

For integration teams weighing servers across the open-source and commercial split, open-source vs commercial FHIR servers for integration stacks is the natural companion read. The benchmark numbers feed into that decision; they do not replace it.