1MIT · 2Cornell University · 3UC San Diego ·
4University of Washington · 5University of Wisconsin–Madison ·
6University of Southern California · 7Purdue University ·
8Futurewei Technologies
* Equal contribution
Ill-defined problems are pervasive in science and engineering, where solutions may be non-existent, non-unique, or lack meaningful interpretation. Despite recent advances of large language models in mathematical reasoning, their ability to reason about ill-defined problems remains poorly understood. Existing benchmarks for unsolvable problems focus on shallow commonsense issues and rely on inflexible evaluation schemes, primarily measuring refusal behavior rather than reasoning capability.
We introduce Ill-Defined Math (IDM), a benchmark of 1,300 ill-defined mathematical problems whose ill-definedness arises from deep, fundamental flaws in their mathematical structure, and a fine-grained three-stage LLM judge that evaluates responses from three complementary perspectives: final-answer statement, issue recognition, and issue fixing. Evaluating 27 recent LLMs reveals substantial degradation compared to well-defined counterparts, and a clear gap between recognizing ill-definedness and handling it correctly: even when models recognize the flaw, they often continue to produce definitive answers without fixing the underlying issue.
The stages run sequentially, with early exit:
Overall correctness is Stage I or Stage II; Stage III is reported separately as a diagnostic. This is what makes "the model saw the flaw and answered anyway" a measurable quantity.
| Model | Final Answer | Recog. | Overall | Well-defined | In-task Abst. Gap |
|---|---|---|---|---|---|
| GPT-5 | 45.2 | 85.3 | 83.6 | 90.7 | 38.5 |
| Gemini 2.5 Pro | 20.0 | 91.7 | 85.3 | 91.3 | 65.3 |
| Grok-4 | 19.1 | 70.1 | 68.4 | 95.9 | 49.4 |
| Claude Sonnet 4 (thinking) | 11.7 | 90.0 | 83.7 | 86.3 | 72.0 |
| o4-mini | 30.0 | 64.3 | 63.0 | 90.0 | 33.0 |
| DeepSeek-V3.1-Think | 9.3 | 96.3 | 88.0 | 94.7 | 78.7 |
| gpt-oss-120b | 26.7 | 92.3 | 90.3 | 87.7 | 63.7 |
| GPT-4o | 18.0 | 28.3 | 27.0 | 61.7 | 9.0 |
| Average over 27 models | 14.7 | 74.0 | 69.1 | 83.1 | 54.4 |
Averaged over 27 models, overall correctness on ill-defined problems is 69.1% while final-answer accuracy is only 14.7%. Models frequently reason their way to a reasonable response — introducing an assumption, repairing a constraint — and then decline to say so in the final answer. Refusal-based metrics therefore substantially underestimate reasoning capability on ill-defined problems, while overall accuracy still falls well short of well-defined performance.
The release contains the 300-problem test set (problem, ill-definedness category, annotation of why it is ill-defined, and the target behavior), the three-stage judge with its prompt templates and evaluation harness, the multi-agent construction pipeline, and the training corpus with its distilled long chain-of-thought supervision.
Test set (300 problems), on Hugging Face:
Code — the three-stage judge, its prompt templates, the evaluation harness, and the multi-agent construction pipeline:
git clone https://github.com/IDMath/IDM.git
@inproceedings{chen2026idm,
title = {Ill-Defined Math: Benchmarking {LLM} Reasoning Beyond Well-Defined Problems},
author = {Chen, Huaibo and Lin, Yixiao and Zhao, Zihan and Chen, Pengcheng and
Liu, Nuohao and Hu, Yue and Xie, Qian and Bai, Qinbo and Yan, Ning and
Mortazavi, Masood S. and Youcef-Toumi, Kamal},
booktitle = {Conference on Language Modeling (COLM)},
year = {2026},
url = {https://openreview.net/forum?id=qgZtkgTwrJ}
}