Applied AI safety research for medicine
Applied AI for better health outcomes, everywhere.
We build AI for health and study how medical AI fails, and who it fails: whether it treats people fairly, stays accurate, and runs privately in a clinic.
dataset downloads, all time, on Hugging Face
About us
AI is entering medicine, and its safety decides whether it helps or harms.
Our work
Datasets
Datasets for clinical AI research, published on Hugging Face.
ClearWrist: Pediatric Wrist X-Ray
20,327 labeled pediatric wrist radiographs, rebuilt from GRAZPEDWRI-DX with clean patient-level splits, verified fracture labels, and bounding boxes.
NEISS Injury Data
7.3 million emergency department injury records from 2005 to 2024, consolidated into a single query-ready file.
NYC Clinic AI Infrastructure
A dataset and interactive map of AI deployment readiness across 637 NYC community health clinic records, spanning 21 languages.
ACNE04
1,457 facial photographs with 4-level severity grades and 18,983 lesion bounding boxes, rebuilt from the original ACNE04 release.
Claude Fable Derm
Patient dermatology questions, each framed with one of six basic emotions, with the model's raw, unprompted answers.
BenchBase MedQA
USMLE-style four-option questions, in the BenchBase format.
BenchBase MedMCQA
187,005 questions from Indian medical entrance exams, in the BenchBase format.
BenchBase PubMedQA
1,000 yes, no or maybe questions with the abstract as context, in the BenchBase format.
BenchBase MMLU Medical
1,242 questions across six medical and biology subjects, in the BenchBase format.
BenchBase TRIAGE
86 mass-casualty scenarios: which triage zone, Red, Yellow, Green or Black.
NYS Health Flyer Repository
A centralized collection of public health materials, organized by language.
Coming soon
Open-source tooling
Tools for testing medical AI the same way every time.
BenchBase
Medical multiple-choice benchmarks in one format. Every model gets the same prompt, every answer is saved, and reports and paired model comparisons are built for you.
Products
Applications built on open-source models.
Research
Papers and experiments on fairness, accuracy and access in clinical AI.
When Education Shouldn't Matter: Counterfactual Bias in LLM-Based Emergency Triage
We tested Qwen-2.5-72B and GPT-4o-mini on 87 clinical vignettes with education-level cues added, holding all medical information constant and measuring how often the decision flipped.
Accepted, ICLR AIMS WorkshopLost in Dialect: Bengali Translation Gaps in NYC Public Health Flyers
Does the Bengali in NYC public health flyers match the dialect its residents understand? We are cataloging flyers by language and using AI to assess dialect accuracy.
In progressDoes Structure Affect Accuracy? Pydantic vs. Unstructured Output on Clinical QA
We compared Pydantic-enforced and unstructured output across GPT-4o-mini, Gemini, and Claude on MedQA questions.
Simplifying Orthopedic Patient Education with Open-Source LLMs
We evaluated open and closed-source models on rewriting OrthoInfo content to an 8th-grade reading level, scored with BERTScore and Flesch-Kincaid grade.