Multispecialty Dental EMRs from Chairside Audio: An Exploratory Study.
Wei L L, Wang Y Y, Lu J J, Wang H H et al.
The objective of this study was to develop and internally evaluate a modular large language model (LLM) system for generating standardized electronic medical records (EMRs) from dental chairside consultations under conditions of acoustic interference and specialty-specific heterogeneity. We built a controllable pipeline integrating multistage audio enhancement and local automatic speech recognition with a cascaded LLM generator. A baseline end-to-end system (system 1) was compared with an evidence-enhanced system (system 2) that added module A (multisource evidence capture and fact table construction) and module B (multiversion collaboration and consensus voting; N = 3, consensus ≥2/3). We evaluated 100 deidentified outpatient recordings from periodontics, orthodontics, and prosthodontics. Output quality was assessed by double-blind clinician scoring and an automated scoring system across completeness, language quality, medical accuracy, and structural standardization (0 to 100 per dimension). System 1 achieved clinically acceptable performance across models, with Kimi-K2 showing the highest stability. System 2 improved robustness with standard deviation reductions of 33.8% (periodontics), 55.9% (orthodontics), and 42.6% (prosthodontics), with the largest gains in complex prosthodontic cases. Analysis of dataset characteristics and acoustic patterns confirmed the clinical authenticity of the corpus and established its suitability for performance evaluation. Artificial intelligence-generated scores correlated with human expert ratings (r = 0.823; R2 = 0.678) with a stable positive bias (mean difference 12.54 points). Our evidence-grounded and consensus-constrained LLM system demonstrated methodological feasibility for generating EMRs from authentic chairside audio, with improvements in medical accuracy and consistency and reduced variability in complex clinical workflows.