COLM 2026

Ill-Defined Math (IDM) Benchmarking LLM Reasoning Beyond Well-Defined Problems

Huaibo Chen1* · Yixiao Lin2* · Zihan Zhao3* · Pengcheng Chen4 · Nuohao Liu5 · Yue Hu6 · Qian Xie2 · Qinbo Bai7 · Ning Yan8 · Masood S. Mortazavi8 · Kamal Youcef-Toumi1

1MIT · 2Cornell University · 3UC San Diego · 4University of Washington · 5University of Wisconsin–Madison · 6University of Southern California · 7Purdue University · 8Futurewei Technologies
* Equal contribution

Abstract

Ill-defined problems are pervasive in science and engineering, where solutions may be non-existent, non-unique, or lack meaningful interpretation. Despite recent advances of large language models in mathematical reasoning, their ability to reason about ill-defined problems remains poorly understood. Existing benchmarks for unsolvable problems focus on shallow commonsense issues and rely on inflexible evaluation schemes, primarily measuring refusal behavior rather than reasoning capability.

We introduce Ill-Defined Math (IDM), a benchmark of 1,300 ill-defined mathematical problems whose ill-definedness arises from deep, fundamental flaws in their mathematical structure, and a fine-grained three-stage LLM judge that evaluates responses from three complementary perspectives: final-answer statement, issue recognition, and issue fixing. Evaluating 27 recent LLMs reveals substantial degradation compared to well-defined counterparts, and a clear gap between recognizing ill-definedness and handling it correctly: even when models recognize the flaw, they often continue to produce definitive answers without fixing the underlying issue.

At a glance

1,300ill-defined problems
300expert-curated test problems
27LLMs evaluated
54.4%avg. In-task Abstention Gap

What makes IDM different

The three-stage judge

The stages run sequentially, with early exit:

  1. Stage I — Final-Answer Statement. Does the final answer explicitly assert that the problem is ill-defined, unsolvable, or not uniquely determined? A naive, overly conservative refusal does not count: the statement must match the actual ill-posedness.
  2. Stage II — Issue Fixing. Does the response repair the problem in a principled way — introducing an explicit assumption, constraining a free variable, or reformulating it so the solution is well-defined under stated conditions?
  3. Stage III — Issue Recognition. Does the reasoning identify the source of the ill-definedness at all, even if the final answer does not act on it?

Overall correctness is Stage I or Stage II; Stage III is reported separately as a diagnostic. This is what makes "the model saw the flaw and answered anyway" a measurable quantity.

Selected results

Accuracy on the IDM test set (%). Final Answer = insolvability declared in the final answer; Recog. = flagged during reasoning; Overall = Stage I or Stage II. The In-task Abstention Gap is Overall − Final Answer, measured entirely within the ill-defined split. See the paper for all 27 models.
ModelFinal AnswerRecog.OverallWell-definedIn-task Abst. Gap
GPT-545.285.383.690.738.5
Gemini 2.5 Pro20.091.785.391.365.3
Grok-419.170.168.495.949.4
Claude Sonnet 4 (thinking)11.790.083.786.372.0
o4-mini30.064.363.090.033.0
DeepSeek-V3.1-Think9.396.388.094.778.7
gpt-oss-120b26.792.390.387.763.7
GPT-4o18.028.327.061.79.0
Average over 27 models14.774.069.183.154.4

The headline finding

Averaged over 27 models, overall correctness on ill-defined problems is 69.1% while final-answer accuracy is only 14.7%. Models frequently reason their way to a reasonable response — introducing an assumption, repairing a constraint — and then decline to say so in the final answer. Refusal-based metrics therefore substantially underestimate reasoning capability on ill-defined problems, while overall accuracy still falls well short of well-defined performance.

Dataset and code

The release contains the 300-problem test set (problem, ill-definedness category, annotation of why it is ill-defined, and the target behavior), the three-stage judge with its prompt templates and evaluation harness, the multi-agent construction pipeline, and the training corpus with its distilled long chain-of-thought supervision.

Test set (300 problems), on Hugging Face:

Code — the three-stage judge, its prompt templates, the evaluation harness, and the multi-agent construction pipeline:

git clone https://github.com/IDMath/IDM.git
The training corpus and its distilled chain-of-thought supervision are still being staged; links here will be updated as each component lands.

BibTeX

@inproceedings{chen2026idm,
  title     = {Ill-Defined Math: Benchmarking {LLM} Reasoning Beyond Well-Defined Problems},
  author    = {Chen, Huaibo and Lin, Yixiao and Zhao, Zihan and Chen, Pengcheng and
               Liu, Nuohao and Hu, Yue and Xie, Qian and Bai, Qinbo and Yan, Ning and
               Mortazavi, Masood S. and Youcef-Toumi, Kamal},
  booktitle = {Conference on Language Modeling (COLM)},
  year      = {2026},
  url       = {https://openreview.net/forum?id=qgZtkgTwrJ}
}