← all benchmarks

TrueMend on SwiftUI — first real baseline

TRUEMEND
57.4%
NO COMPETITOR
BASELINE · NO FIXES YET

First real baseline, published before any remediation: 646 findings judged on a shipping App Store app, 14.2% real-rate. SwiftLint exists but is not in the harness yet. Prior evidence had tested a networking library with zero UI code; retired the same way as Ktor’s.

First real blind-adjudicated baseline for SwiftUI: 57.4% wrong-match across 646 judged findings on a real shipping App Store app. SwiftLint isn't installed in this harness, so there's no competitor number yet, and no fixes have been applied — the honest, unimproved starting point, published before remediation, not after.

Same discipline as the Ktor baseline published alongside this one: a first measurement, published as-is, before any fix work — not held back until it looks better.

No competitor number, this round

SwiftLint exists as a tool but isn't installed in this harness yet, so this is TrueMend-only. Judged the same way every other campaign is judged — blind, two-axis, against real production source.

The corpus, and a corrected mistake

646 findings, blind two-axis judged, drawn from Dimillian/IceCubesApp — a real, shipping App Store app, 391 files. Prior evidence for this framework tested Alamofire, a networking library with zero UI code in it — the wrong corpus, by mistake, the same way Ktor's prior evidence tested a SQL library instead of a server framework. That result is retired; this is the first evidence that actually exercises SwiftUI-specific detectors against real view code.

Result

57.4% wrong-match, 14.2% real-rate, 646 judged. No fixes have been applied. Round one is measurement, not remediation, per the same campaign discipline applied to Ktor, Django, and every other first-round baseline on this site.

Still open

57.4% is within the methodology's own "expect 20-50% on a first run" range but at the high end of it. Nothing has been fixed yet, and there's no competitor number to compare against until SwiftLint is wired into the harness. Next step is the same as Ktor: root-cause the top wrong categories against real IceCubesApp source, fix detection logic (not just the judged sites), and re-certify on a fresh sample before claiming any improvement.