Patientdesk Labs

The research layer for dental clinic AI

Patientdesk Labs is the research arm of Patientdesk.ai. We build benchmarks, fine-tune models, and develop domain-specific tools so that when an AI answers the phone at a dental office, it actually works.

What We Work On

Four Research Tracks

Dental clinic AI has unique requirements that generic models don't meet. Our research focuses on four areas where domain-specific work makes the biggest difference — from the model that reasons, to the voice the patient actually hears.

Evaluation
Live

DentesBench

The first benchmark for evaluating LLMs as dental clinic phone agents. 483 scenarios across 10 categories, scoring empathy, clinical safety, accuracy, brevity, and tone — plus a deployment-weighted leaderboard.

Read the paper →
Language Models
WIP

Gemma 4 Fine-Tuning

Soul-document-driven fine-tuning of Gemma 4 31B with Opus 4.6 as judge. After one iteration of SFT + DPO: 8.46 on DentesBench — beating every frontier API model including Opus itself.

Read the paper →
Speech Synthesis
Live

AI Voice Benchmark

A blind listening leaderboard for AI voices. Votes come from a vetted panel of paid native speakers rather than whoever finds the page, and each listener answers one of four questions instead of a single “which is better”. Across 12,245 comparisons, Fish Audio’s s2-pro is 5th of 44 on sounds human and 40th on clear and correct — and Cartesia’s sonic-3.5 is 3rd in English, 19th in Turkish.

Open the arena →
Speech Recognition
Future Work

Dental STT

Adapting speech-to-text models for dental clinic phone audio. Patient calls with accents, background noise, and dental terminology that generic models consistently get wrong — "prophylaxis" shouldn't become "prophy lax is."

Paper coming soon
Our Approach

Why Dental AI Needs Its Own Research

A dental receptionist AI has to be warm without accidentally diagnosing, efficient without being cold, and helpful without overstepping clinical boundaries. No off-the-shelf model gets this right consistently.

Measure What Matters

Generic benchmarks test reasoning and knowledge. We test whether a model can be empathetic to an anxious patient without crossing clinical lines.

Train for the Domain

Frontier models know dental terminology. What they lack is the discipline to stay warm and safe simultaneously. That requires domain-specific training.

Respect Patient Privacy

All research data is fully de-identified in compliance with HIPAA regulations. No Protected Health Information is used in any benchmark or training process.

Release 2026.2

AI Voice Arena

A vetted panel of paid native speakers listens to two AI voices reading the same sentence and picks one, without ever seeing who made either. Each sitting asks one of four questions and never mixes them: overall preference, sounds human, clear and correct, rhythm and expression. 12,245 comparisons across 44 English and 28 Turkish voices from 16 companies, judged by 224 listeners and frozen on 25 July. The two boards below both show overall preference, one language each — note how little they share. Open the full arena →

English — overall preference2185 votes · 103 listeners
1 Google GeminiPuck 1138
2 Smallestnolan 1125
3 Cartesiadb6b0ed5 1124
4 Speechifygeffen_32 1121
5 xAIaltair 1107
6 Speechifydominic_32 1100
Turkish — overall preference851 votes · 44 listeners
1 Google GeminiPuck 1302
2 Google GeminiKore 1195
3 Google GeminiKore 1185
4 Google GeminiPuck 1183
5 xAIara 1157
6 MiniMaxTurkish_CalmWoman 1151
95% range score bottom of leader’s range

With 4.6× the data of our pilot, English now separates: 20 of the 43 challengers sit clearly behind the leader. The language gap is still the biggest single finding — Cartesia’s sonic-3.5 is 3rd in English and 19th in Turkish, ElevenLabs’ eleven_v3 is 37th and 8th, Google’s Gemini 2.5 Puck 38th and 4th; Gemini 3.1 Flash Puck, top of both boards above, is the exception. A vetted panel is not a representative one, and it is not infallible either: 9.3% of the hidden same-clip controls still got a confident winner picked.

Open the Arena Read the Report
Latest Results

DentesBench v0.2 Leaderboard

Eight models evaluated on 483 dental phone agent scenarios. The v2 score weights quality (80%), cost (10%), and latency (10%) to reflect real deployment constraints. Full methodology →

Top 5 — Deployment-Weighted Ranking April 2026
#Model EmpathySafety AccuracyBrevity ToneV2 Pass Latency Cost/resp
1Gemma 4 31B (OpenRouter) 6.89.6 7.38.7 6.88.18 75% 2.4s $0.00006
2GLM-5 Turbo (OpenRouter) 7.09.7 7.78.7 7.38.01 84% 3.5s $0.00074
3GPT-5.4 6.99.6 8.08.3 7.07.92 86% 1.5s $0.00161
4Claude Sonnet 4.6 7.49.5 7.78.3 7.77.86 88% 2.3s $0.00189
5Claude Opus 4.6 7.59.6 7.88.3 7.87.41 91% 3.1s $0.00318

Interested in Our Research?

Read the full DentesBench paper for methodology, results, and what we've learned about the tradeoffs in dental AI.

Read the Paper Visit Patientdesk.ai