Research

I work on measurement and systems problems where a plausible abstraction can quietly stop matching the thing underneath it.

My standard is evidence-first: primary artifacts before prose, negative results remain negative, targets remain targets, and a later validity failure can weaken an earlier headline rather than being written around. The correction history is public.

AI evaluations · security

BRIDGE-bench — when the benchmark leaks its own answer

BRIDGE-bench evaluates model reasoning over verified smart-contract exploit examples. A later prompt audit found severe author-injected leakage in 13 of 24 examples. In the sharpest Euler case, the prompt header contained the exact ground-truth label while the vulnerable function was absent from the committed source bundle.

The historical model score is therefore a contaminated upper bound, not clean vulnerability-detection performance. The repository now has comment-stripping, identity anonymization, sanitizer invariants, and a locked paired remeasurement design. I am not claiming the sanitized score because the paid inference rerun is intentionally not part of the current program.

Scientific ML

Active materials discovery — useful mean, failed uncertainty

On a static 18,928-structure matbench_perovskites screening task, the pretrained mean prediction reaches roughly 5× discovery acceleration relative to random at the tested 5% budget. A low-weight UCB rule ties greedy ranking.

Direct calibration changes the interpretation: the readout-level MC-dropout standard deviation has Spearman −0.47 correlation with absolute error, and increasing its acquisition weight eventually pushes DAF below random. The current result is not “uncertainty helps”; it is that a good point predictor and a good uncertainty estimator are separate achievements.

Systems · networking

roce-preflight — independent state as a test boundary

The project had 197 passing unit tests before the first real RDMA path exposed four defects in MTU reporting, GID diagnosis, route selection, and GID enumeration. The first attempted fix passed the unit suite and still failed on the device.

The result is a test architecture, not a benchmark brag: logic stays fast and synthetic; claims about Linux RDMA require real verbs and sysfs state; physical-NIC/fabric performance remains a separate evidence class.

Autonomy · hardware

Aiur — prototype the risky interaction, not the rendering

Aiur is a persistent airborne-carrier concept. CARRIER-P0 removes hydrogen, swarms, charging, outdoor autonomy, and heavy airborne compute until one falsifiable interaction remains: repeated mechanically verified recovery of a small aircraft onto a moving dock.

The current artifacts are design, models, CAD, controller logic, and acceptance gates. Recovery rates are targets, not observed results, until a physical telemetry dataset is committed.

Method

I use four evidence classes in technical work: observed, derived, estimated, and target. The article should never make one look like another. The strongest rival explanation belongs in the main argument, not a footnote.

Public revisions → · Formation → · engineering thesis →