Under review at DMLR · dataset releasedDMLR (Journal of Data-centric Machine Learning Research)·Benchmark · dataset
Quantifying Chart Deception: Continuous Deception-Severity Estimation for Chart-Reading Models
Carson RodriguesiD (Celabe)
Charts can mislead without containing a single false number: a truncated axis, an inverted scale, or a stretched aspect ratio changes what a reader concludes while every label stays accurate. This paper makes deception a continuous, exactly computable quantity: a counterfactually-augmented dataset generated directly from distortion parameters (2,430 ground-truth rows carrying a 33-field machine-checkable schema, 636 with an exact Tufte lie factor, plus a 2,129-row failure-weighted training projection), a base-model difficulty map showing accuracy collapses exactly where the deception lives (continuous severity, extraction under a distorted axis, mechanism naming), and a held-out Deception Atlas of 20 real published misleading charts verified against primary sources. Training against the exact ground truth with reinforcement learning from verifiable rewards is pre-committed in an analysis plan fixed before any run.
2,430 exact-ground-truth rows · 636 exact lie factors · 20-chart Deception Atlas · RLVR plan fixed in advanceChart understandingVision-language modelsRLVRBenchmarksShortcut learning
Sole-authored. The dataset, 954 chart renders, and the Deception Atlas are released under CC BY 4.0 (Zenodo DOI 10.5281/zenodo.21939803); the severity target is computed from each chart's construction, so the reward needs no human rater and no judge model. The RLVR training runs follow the fixed plan and will appear in a follow-up paper. Preprint published and submitted to DMLR on 15 August 2026. Companion to Paper 15.