Applied AI safety research for medicine

Applied AI for better health outcomes, everywhere.

We build AI for health and study how medical AI fails, and who it fails: whether it treats people fairly, stays accurate, and runs privately in a clinic.

dataset downloads, all time, on Hugging Face

About us

AI is entering medicine, and its safety decides whether it helps or harms.

Our work

Datasets

Datasets for clinical AI research, published on Hugging Face.

  • ClearWrist: Pediatric Wrist X-Ray

    20,327 labeled pediatric wrist radiographs, rebuilt from GRAZPEDWRI-DX with clean patient-level splits, verified fracture labels, and bounding boxes.

  • NEISS Injury Data

    7.3 million emergency department injury records from 2005 to 2024, consolidated into a single query-ready file.

  • NYC Clinic AI Infrastructure

    A dataset and interactive map of AI deployment readiness across 637 NYC community health clinic records, spanning 21 languages.

  • ACNE04

    1,457 facial photographs with 4-level severity grades and 18,983 lesion bounding boxes, rebuilt from the original ACNE04 release.

  • Claude Fable Derm

    Patient dermatology questions, each framed with one of six basic emotions, with the model's raw, unprompted answers.

  • BenchBase MedQA

    USMLE-style four-option questions, in the BenchBase format.

  • BenchBase MedMCQA

    187,005 questions from Indian medical entrance exams, in the BenchBase format.

  • BenchBase PubMedQA

    1,000 yes, no or maybe questions with the abstract as context, in the BenchBase format.

  • BenchBase MMLU Medical

    1,242 questions across six medical and biology subjects, in the BenchBase format.

  • BenchBase TRIAGE

    86 mass-casualty scenarios: which triage zone, Red, Yellow, Green or Black.

  • NYS Health Flyer Repository

    A centralized collection of public health materials, organized by language.

    Coming soon

Open-source tooling

Tools for testing medical AI the same way every time.

  • BenchBase

    Medical multiple-choice benchmarks in one format. Every model gets the same prompt, every answer is saved, and reports and paired model comparisons are built for you.

Products

Applications built on open-source models.

  • Wellspring

    Bank your voice while you still have it. Record a few sentences now, and later generate speech in your own voice from typed text. It runs on your own device.

  • GroveAI

    Free, private AI for everyone.

  • Hue

    Quantify and track facial skin texture and redness over time.

Research

Papers and experiments on fairness, accuracy and access in clinical AI.

  • When Education Shouldn't Matter: Counterfactual Bias in LLM-Based Emergency Triage

    We tested Qwen-2.5-72B and GPT-4o-mini on 87 clinical vignettes with education-level cues added, holding all medical information constant and measuring how often the decision flipped.

    Accepted, ICLR AIMS Workshop
  • Lost in Dialect: Bengali Translation Gaps in NYC Public Health Flyers

    Does the Bengali in NYC public health flyers match the dialect its residents understand? We are cataloging flyers by language and using AI to assess dialect accuracy.

    In progress
  • Does Structure Affect Accuracy? Pydantic vs. Unstructured Output on Clinical QA

    We compared Pydantic-enforced and unstructured output across GPT-4o-mini, Gemini, and Claude on MedQA questions.

  • Simplifying Orthopedic Patient Education with Open-Source LLMs

    We evaluated open and closed-source models on rewriting OrthoInfo content to an 8th-grade reading level, scored with BERTScore and Flesch-Kincaid grade.

If you love how AI could be used safely for medicine, join us.

Join us