Wenjing Yue

Hallu-TCM

We develope a novel TCM hallucination detection dataset, Hallu-TCM, sine no prior work has attempted this task in TM. We selected 1,260 TCM exam questions including 16 TCM subjects, input them into GPT-4, and collected their feedback. In the first level, we utilize Qwen-Max interface to annotate feedback multiple times with the binary label. If Qwen-Max consistently provided the same label across annotations, we adopted that label. For contentious cases, we recruited higher-degree research students who can understand and solve complex questions, including three Ph.D.

Categories:

Artificial Intelligence
Education and Learning Technologies
Reliability
Biomedical and Health Sciences

Dataset Entries from this Author

Hallu-TCM