benchmark

We publish the whole benchmark, including where we lose.

A security tool that raises too many false alarms gets ignored by the team. Here is codafort measured against CodeQL and Semgrep on our benchmark tests, with the caveats next to the numbers.

Java: all three tools in the same run

Our benchmark tests, with a known answer for every case, measured on 2026-09-28. Precision: how many alerts were real. Recall: how many real flaws were found.

toolprecisionrecallfalse positives
codafort1.0000.9720
CodeQL 2.270.7210.972531
Semgrep OSS 1.1770.6930.882552
0 vs 531
false positives against CodeQL, finding the same real flaws
0.955
precision on Python, at 0.947 recall; CodeQL scored 0.432 and Semgrep 0.828
8 s
to analyze the Java tests, against 15 to 20 s for Semgrep and 99 s for CodeQL
⚠ The caveat comes before the number. We tuned codafort looking at the Java tests; CodeQL and Semgrep did not. Treat the result as a ceiling, not a guarantee: on real code, every analyzer has some false positives. On 349 cases across 16 languages that no tuning could use, precision is 1.000 and recall 0.762. On the tests for the other languages, below, the ranking holds and the gaps shrink.

Per language, against CodeQL and Semgrep

Precision · recall on our benchmark tests for each language, in the same run (2026-09-28). For Semgrep, its best configuration.

languagecodafortCodeQLSemgrep
Java1.000 · 0.9720.721 · 0.9720.693 · 0.882
Python0.955 · 0.9470.432 · 0.2520.828 · 0.159
JavaScript0.943 · 0.8110.877 · 0.7790.756 · 0.239
TypeScript0.931 · 0.900—¹0.714 · 0.167
C#1.000 · 0.941—²0 hits
Go0.919 · 0.358—²0.852 · 0.164
Ruby1.000 · 0.1920.707 · 0.0820.755 · 0.105
Rust1.000 · 0.920—¹1.000 · 0.080
Swift1.000 · 0.939—¹0 hits
PHP1.000 · 0.867n/a1.000 · 0.333
C0.968 · 0.3950.667 · 0.0260 hits
C++0.800 · 0.0980 hits0 hits

¹ A test written by CodeQL itself: we do not measure CodeQL on it. ² Not measured in this run. n/a: CodeQL does not analyze PHP. On the full NIST Juliet for Java, codafort has precision 0.88 and recall 0.81 (2026-10-02).

Where we still lose

On precision, to CodeQL, on TypeScript from real CVEs (720 cases): 0.831 for CodeQL against 0.796 for us, at similar recall.

On recall in Go, Ruby, Rust, C and C++. We are still ahead on these tests, but far from 0.90, and the number is low for every tool measured.

Every tool was measured in the same run, by the same criteria, on our benchmark tests.

Precision that is measured, published and compared.

Pre-launch: codafort is not available to install yet. Join the waitlist →