research integrity

Revisions

A result is allowed to become less flattering when better evidence arrives.

This page is generated from the same audit ledger that labels every article on the site. I keep corrections visible because the ability to weaken your own claim is part of technical judgment, not an embarrassment to hide in Git history.

revised research note2026-08-08

When uncertainty makes active learning worse

The original version treated Greedy≈UCB as evidence that the uncertainty signal was useful or the pretrained surrogate was well calibrated. Direct calibration later found readout-level MC-dropout σ anti-correlated with absolute error (Spearman −0.47) on the current perovskite benchmark. The article is revised in place: the mean prediction is useful; this uncertainty estimate is not under the tested configuration.

Primary evidence: active-materials-discovery ↗

revised evaluation postmortem2026-08-08

A benchmark can measure its own metadata

The original ~40% LLM F1 headline is not clean capability evidence. A later audit found severe author-injected prompt leakage in 13/24 verified contracts, including an Euler example whose header contained the exact label while the vulnerable function was absent from committed source. The article is revised in place and the historical score is treated as a contaminated upper bound; no paid sanitized rerun is claimed.

Primary evidence: BRIDGE-bench source + leakage audit ↗

merged upstream2026-08-08

Recursive types, finite values: an EIP-712 bug in alloy

Revised after Alloy #1105 merged. The current treatment distinguishes legal self-reference during EIP-712 canonical type encoding from strict runtime resolution, which still cannot represent recursive values; mutual recursion remains rejected.

Primary evidence: Alloy PR ↗

revised security note2026-08-08

An invariant is not an end-to-end proof

Revised to make the trust boundary explicit: supply≤collateral is a local accounting invariant, while an attestation gate makes an operator-coordination assumption enforceable on-chain. A signature proves who authorized a statement; it does not independently prove a remote-chain event.

Primary evidence: lock-mint-bridge-lab ↗

design note2026-08-08

Delete the science-fiction parts first

This is deliberately a design note, not a test report. Aiur has specifications, models, CAD, a dock controller, and acceptance gates, but no committed physical recovery dataset yet. Recovery rates and closing-speed limits remain targets until instrumented telemetry exists.

Primary evidence: Aiur CARRIER-P0 ↗

Policy

Observed, derived, estimated, and target are different evidence classes. A later audit can downgrade an earlier headline. A design target never becomes a measured result because the prose is persuasive.

Publication standard → · research programs →