1st place · Data & Visualization · $4,000

Adaption Labs · AutoScientist Challenge, Part 2

PolyChart — "Shown Is Not Supported"

One model that scores how much a chart lies, submitted to all seven Part 2 categories. A chart can be true in every value and still misrepresent its result. PolyChart measures how much a chart distorts its data, giving every distortion an exact Tufte Lie Factor, signed and continuous, rather than a yes/no label. Everything below is open: datasets and weights on HuggingFace and Kaggle, live demos, the client that drove the whole run, and the research it produced.

7
Categories entered
85.9
Best held-out
176
Fine-tune jobs
13
Datasets

The result

$4,0001st place in Data & Visualizationannounced 18 August 2026
A visualization-reasoning model designed to determine whether claims made from charts are actually supported by the underlying data and to detect misleading visual patterns.
Adaption Labs, announcing the AutoScientist Challenge, Part 2 winners

Part 2 ran seven categories — Science, Agriculture, Personal Finance, Data & Visualization, Market Analysis & News, HR, and Math & Code — with a first and a runner-up named in each. The runner-up in this category was Vinod Anbalagan. The entry below is the one that placed; the other six shared the same model and the same thesis.

The automation that made it possible

To run at the scale this needed — 176 fine-tuning jobs across 86 launches and 13 datasets — I built a headless client for the AutoScientist API before the official one shipped, and drove everything through it rather than the interface. It is what made the paired within-seed design and the run count feasible. I have open-sourced it so anyone can build on it.

The seven submissions

6 July – 10 August 2026

One idea, tested seven ways. Each category is its own dataset and training run, evaluated on the platform's held-out test set for that category. All seven share the single LoRA released below.

Data Visualization

1st place
85.9held-out

A model that scores how much a chart lies. Every distortion carries an exact Tufte Lie Factor, signed and continuous, not a yes/no label. Ships with a Deception Atlas of 21 real published misleading charts, each value checked against primary sources.

dataset de9610d4-9df5-40a9-abee-b95909d7d736model ccbc149b-5943-494b-8d13-91c95329dc93

Science

83.9held-out

Scientific claims travel as charts, and a chart can be true in every value and still misrepresent its result. This checks the claim against the plotted data, not the picture.

dataset 2f5a3428-1be6-4979-b969-cc1fe19159d1model 7f4c4763-d888-44eb-a1ba-696749529b2a

Market Analysis & News

80.5held-out

News charts are the most published and least audited charts there are. This pairs chart reading with severity scoring, so a claim in a story can be checked against the figure that supposedly supports it.

dataset 59389e67-4027-46db-b98c-f46ec2e27f07model 7ab73987-64ff-41d3-bc9e-26f593257115

HR

77.0held-out

Workforce reporting is full of traps: a comparison across differently defined cohorts, an axis choice that carries a conclusion the data does not. It reads participation and trend claims against the underlying numbers.

dataset 4b2a3093-6313-4130-9b89-4ddfdc291c72model bd802f46-0562-4bf8-89a6-921324753408

Math & Code

77.0held-out

Quantitative reasoning over tabular and plotted data. Read exact values, compute ratios and rates of change, and refuse to answer when the number is not in the input. The refusal case matters as much as the arithmetic.

dataset 4c5bb88a-653b-4415-8f17-4dc0261fe217model 6431e17e-19bc-49ad-8f29-7dca70744f39

Personal Finance

76.5held-out

Financial charts are where misleading encodings do the most direct damage: a truncated axis on a returns chart, a percentage gain quoted without its base. It separates the visual impression from the arithmetic and reports the size of the gap.

dataset 394b0591-7c2b-415f-a314-2bbec775700emodel 7c98d08c-170c-447e-8ab3-99091c5022c2

Agriculture

75.5held-out

Agricultural yield and input data, where long series and unit changes make trend claims easy to overstate. Measure first, then claim only what the data supports.

dataset b2e55669-d91c-4ad7-82c9-2cadc6c7caeamodel acc81241-3a67-4c51-8668-d7a363ba69a1

Shared model weights

A single LoRA, PolyChart — "Shown Is Not Supported", released for every category on both platforms.

Research

Two papers came out of this build. Neither is on arXiv yet; both are honest work in progress, and I will post the preprints here and in the Adaption community when they are up. The fuller list lives on my research page.

Under review at DMLR · Zenodo preprint

Silent Failure in Automated Model Adaptation

The measurement paper from this build. It instruments the platform across 176 fine-tuning jobs, 86 launches, and 13 datasets, driven headlessly through its internal API. Two findings, usually reported apart: the core capability works (synthetic augmentation improved the held-out score on 11 of 11 paired seeds, mean +14.4, 95% bootstrap CI 12.4 to 16.4, sign test and Wilcoxon p=0.0010, and the advertised +16 from crossing 20k datapoints reproduced at +17.4), and the reported numbers still cannot, on their own, separate a good adaptation from a bad one, with the displayed and held-out metrics disagreeing in sign in three of eight paired comparisons and the quality grade collapsing toward a fixed attractor. It defines silent adaptation failure, proposes six cheap machine-checkable audits, and reports a verifier that caught nine failures in the pipeline, four in its own scaffolding.

176 fine-tunes · +14.4 on 11/11 seeds (CI 12.4–16.4, p=0.0010) · 6 audits
Under review at DMLR · dataset released

Quantifying Chart Deception: Continuous Deception-Severity Estimation for Chart-Reading Models

The model paper behind PolyChart. Charts can mislead without a single false number: a truncated axis or a mismatched area encoding changes what a reader concludes while every label stays accurate. It builds a counterfactually-augmented dataset directly from the distortion parameters used to render each chart, so the ground truth for deception is generated rather than annotated — 2,430 rows on a 33-field machine-checkable schema, 636 carrying an exact Tufte lie factor — plus a held-out Deception Atlas of 20 real published misleading charts verified against primary sources. Training a continuous severity estimator with reinforcement learning from verifiable rewards is pre-committed in an analysis plan fixed before any run.

2,430 exact-ground-truth rows · 636 exact lie factors · 20-chart Deception Atlas

Available for senior AI / contract / FDE work

Building something with AI?

Voice agents, MCP servers, LLM pipelines, agentic workflows — pick a slot, drop a message, or send your email and I'll reply within a day.

or leave your email

Replies within ~24 hours · Remote-first · global · open to relocation