Compare LLM training paradigms on ESI triage classification (MIETIC-36 expert set):
โ ๏ธ Research demo only โ not for clinical use. Always consult a qualified clinician.
SFT = supervised fine-tuning ยท DPO = preference optimization ยท GRPO = RL on ESI rewards
Show model's step-by-step reasoning (Qwen3 extended thinking)
GRPO + thinking needs โฅ1024 to fit reasoning + EXTRACTION/ALGORITHM/ANSWER