Large Language Models offer transformative potential for personalized learning, yet they frequently fail at pedagogical scaffolding-leaking answers or using circular reasoning when generating hints. We present a framework that orchestrates human and AI agency to produce high-fidelity multilingual Socratic hints at scale. Human annotators define quality standards and calibrate automated critics; autonomous AI agents then enforce these standards through adversarial generation, critique, refinement, and recovery loops. This division of agency-humans set the rubric, AI operates at scale-produces MultiHint-9K: a dataset of 9,987 machine-verified, human-calibrated QA-Hint tuples across English, Italian, and Farsi, with a 97.5% yield rate. The pipeline's adversarial critics flagged 5,743 errors across generation attempts, including 3,131 circular reasoning and 495 answer leakage instances. We fine-tune five open-source models (7B-24B) via LoRA, achieving ∼67% ROUGE-L improvement over base models (all p < 0.001), with the largest gains in low-resource Farsi (∼97% ROUGE-L). A component contribution analysis demonstrates that removing any pipeline stage degrades yield from 98% to as low as 38%. We publicly release the dataset, models, and codebase at https://github.com/KamyarZeinalipour/MultiHint-9K.

Zeinalipour, K., Sadeghi, A., Angelini, G., Rigutini, L., Gori, M., Maggini, M. (2026). MultiHint-9K: Orchestrating Human and AI Agency for Scalable Pedagogical Hint Generation. In CEUR Workshop Proceedings (pp.144-150). CEUR-WS.

MultiHint-9K: Orchestrating Human and AI Agency for Scalable Pedagogical Hint Generation

Gori, M.;Maggini, M.
2026-01-01

Abstract

Large Language Models offer transformative potential for personalized learning, yet they frequently fail at pedagogical scaffolding-leaking answers or using circular reasoning when generating hints. We present a framework that orchestrates human and AI agency to produce high-fidelity multilingual Socratic hints at scale. Human annotators define quality standards and calibrate automated critics; autonomous AI agents then enforce these standards through adversarial generation, critique, refinement, and recovery loops. This division of agency-humans set the rubric, AI operates at scale-produces MultiHint-9K: a dataset of 9,987 machine-verified, human-calibrated QA-Hint tuples across English, Italian, and Farsi, with a 97.5% yield rate. The pipeline's adversarial critics flagged 5,743 errors across generation attempts, including 3,131 circular reasoning and 495 answer leakage instances. We fine-tune five open-source models (7B-24B) via LoRA, achieving ∼67% ROUGE-L improvement over base models (all p < 0.001), with the largest gains in low-resource Farsi (∼97% ROUGE-L). A component contribution analysis demonstrates that removing any pipeline stage degrades yield from 98% to as low as 38%. We publicly release the dataset, models, and codebase at https://github.com/KamyarZeinalipour/MultiHint-9K.
2026
Zeinalipour, K., Sadeghi, A., Angelini, G., Rigutini, L., Gori, M., Maggini, M. (2026). MultiHint-9K: Orchestrating Human and AI Agency for Scalable Pedagogical Hint Generation. In CEUR Workshop Proceedings (pp.144-150). CEUR-WS.
File in questo prodotto:
File Dimensione Formato  
short11.pdf

accesso aperto

Tipologia: PDF editoriale
Licenza: Creative commons
Dimensione 1.27 MB
Formato Adobe PDF
1.27 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11365/1326834