Large Language Models offer transformative potential for personalized learning, yet they frequently fail at pedagogical scaffolding-leaking answers or using circular reasoning when generating hints. We present a framework that orchestrates human and AI agency to produce high-fidelity multilingual Socratic hints at scale. Human annotators define quality standards and calibrate automated critics; autonomous AI agents then enforce these standards through adversarial generation, critique, refinement, and recovery loops. This division of agency-humans set the rubric, AI operates at scale-produces MultiHint-9K: a dataset of 9,987 machine-verified, human-calibrated QA-Hint tuples across English, Italian, and Farsi, with a 97.5% yield rate. The pipeline's adversarial critics flagged 5,743 errors across generation attempts, including 3,131 circular reasoning and 495 answer leakage instances. We fine-tune five open-source models (7B-24B) via LoRA, achieving ∼67% ROUGE-L improvement over base models (all p < 0.001), with the largest gains in low-resource Farsi (∼97% ROUGE-L). A component contribution analysis demonstrates that removing any pipeline stage degrades yield from 98% to as low as 38%. We publicly release the dataset, models, and codebase at https://github.com/KamyarZeinalipour/MultiHint-9K.
Zeinalipour, K., Sadeghi, A., Angelini, G., Rigutini, L., Gori, M., Maggini, M. (2026). MultiHint-9K: Orchestrating Human and AI Agency for Scalable Pedagogical Hint Generation. In CEUR Workshop Proceedings (pp.144-150). CEUR-WS.
MultiHint-9K: Orchestrating Human and AI Agency for Scalable Pedagogical Hint Generation
Gori, M.;Maggini, M.
2026-01-01
Abstract
Large Language Models offer transformative potential for personalized learning, yet they frequently fail at pedagogical scaffolding-leaking answers or using circular reasoning when generating hints. We present a framework that orchestrates human and AI agency to produce high-fidelity multilingual Socratic hints at scale. Human annotators define quality standards and calibrate automated critics; autonomous AI agents then enforce these standards through adversarial generation, critique, refinement, and recovery loops. This division of agency-humans set the rubric, AI operates at scale-produces MultiHint-9K: a dataset of 9,987 machine-verified, human-calibrated QA-Hint tuples across English, Italian, and Farsi, with a 97.5% yield rate. The pipeline's adversarial critics flagged 5,743 errors across generation attempts, including 3,131 circular reasoning and 495 answer leakage instances. We fine-tune five open-source models (7B-24B) via LoRA, achieving ∼67% ROUGE-L improvement over base models (all p < 0.001), with the largest gains in low-resource Farsi (∼97% ROUGE-L). A component contribution analysis demonstrates that removing any pipeline stage degrades yield from 98% to as low as 38%. We publicly release the dataset, models, and codebase at https://github.com/KamyarZeinalipour/MultiHint-9K.| File | Dimensione | Formato | |
|---|---|---|---|
|
short11.pdf
accesso aperto
Tipologia:
PDF editoriale
Licenza:
Creative commons
Dimensione
1.27 MB
Formato
Adobe PDF
|
1.27 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11365/1326834
