Benchmark complete · expert-validated · AAAI-27 build readyAAAI 2027 (target)·Conference
Done Is Not Correct: Measuring Silent Failures and Self-Verification Calibration When LLM Agents Take CAD Actions
Carson RodriguesiD (Celabe), Clive Rodrigues, Aravind Reddy G
A reproducible benchmark for agentic CAD. When an LLM agent writes parametric CAD code from a natural-language spec, syntactic success (the code runs, the geometry renders) decouples from semantic correctness (right dimensions, valid constraints, manufacturable features). The harness measures that silent-failure gap directly and tests whether claim-conditioned self-verification closes it.
468 runs · single-shot 14.1% silent failures (ECE 0.32) → 3.2% with enforced self-checks · expert–grader agreement κ = 0.64 (0.87 excluding one ambiguous task)Agentic safetyReliabilityEvaluation
Benchmark complete (4 models × 3 conditions); independent CAD-engineer validation done (blind pass/fail scoring of 81 parts). Two-column AAAI-27 anonymous manuscript built; submitting to AAAI 2027 (abstracts due 21 Jul 2026).