Radar · 01/08/2026 · happened on 31/07/2026 · models

GPT-5.6 Disproves Maxwell's Conjecture, Distillation Copies Abilities Without Censorship

What happened

A paper on arXiv (July 29, 2026) uses GPT-5.6 to find a counterexample to Maxwell’s conjecture, an open problem in electrostatics since 1864. Five point charges in Euclidean space with at least 24 non-degenerate critical points: beyond the $(n-1)^2$ limit that Maxwell had proposed. The frontier model helped close a problem mathematicians hadn’t solved.

Simultaneously, CTGT publishes research on distillation. Training GPT-OSS-120B on DeepSeek V4 Flash outputs for financial reasoning, the capabilities transfer, the censorship doesn’t. The distilled model describes Uyghur transfer programs and the Great Leap Forward without any political reformulation. Across 152 controlled prompt pairs and four judges from different labs, no statistically significant difference from the untouched base model.

Why it matters to you

The frontier still solves problems open models don’t touch. The distillation pipeline means Chinese open models’ capabilities can be absorbed without dragging along their political constraints. As we reported on August 1st, DeepSeek V4-Flash already matches GPT-5.6 Luna at 60% lower task cost: the gap between open and proprietary shrinks from both sides. Open models approach on raw performance, and their useful behaviors transfer clean.

In detail

Maxwell’s Conjecture

Maxwell’s conjecture (1864) concerns the electrostatic field of $n$ point charges. Maxwell hypothesized the potential has at most $(n-1)^2$ critical points, all non-degenerate. For $n = 5$, the limit would be 16. Philip Arathoon, Gavin Ball and Matthew D. Kvalheim exhibit a configuration of five charges with at least 24 non-degenerate critical points. The conjecture is false.

The paper, posted to arXiv on July 29, 2026, identifies GPT-5.6 as a tool in the discovery or verification process of the counterexample. It signals frontier models enter as exploratory tools in pure research, beyond code writing and data analysis. The caveat: a counterexample found with model help should be verified with formal methods, and the paper is a preprint (peer review still underway).

CTGT’s experiment

Johnny Yu, Siddarth Mamidanna and Cyril Gorlla tested whether censorious behavior transfers alongside capabilities, using DeepSeek V4 Flash as teacher and GPT-OSS-120B as student on financial reasoning.

DeepSeek V4 Flash showed a 45.45 point gap on politically sensitive prompts versus structurally identical controls (152 pairs, four judges from four frontier labs). The distilled model showed no statistically significant difference from the base model. It described Uyghur transfer programs, the Great Leap Forward, Tiananmen, without reformulation.

Self-distillation (the model corrects its own errors and trains itself) matched distillation from the Chinese model on every seed. The distilled model scores 83.61% on FinanceReasoning, above Kimi K3 (81.93%) and Inkling (65.13%), at 62x lower query cost than Inkling and 160x lower than Kimi K3.

CTGT released everything: 304 prompts, paired controls, judges’ rubric, evaluation code, the models themselves.

Limits and implications

The result is specific to finance domain and Chinese censorship. It doesn’t prove every unwanted behavior filters away through distillation, nor that every capability transfers clean. The sample is moderate. CTGT has commercial interest in the narrative that Chinese models are safe to distill: data are public and verifiable, but generalization warrants caution.

The self-distillation finding may be most relevant for builders. If confirmed across domains, a model can achieve the same gains by correcting its own errors, without an external teacher. It would further reduce larger teacher model advantage and close another piece of the gap between those with frontier access and those without.

Type to search across course, playbooks, skills, papers…