Measuring the Sloppiness of Code: Verbosity and Erosion Metrics
LLMs have gotten very good at generating code that passes tests. But correct code can still be sloppy: full of duplicated lines, unnecessary abstractions, and poor design choices. As the author notes, "just because the code is formally correct doesn't mean that it is not introducing unnecessary abstractions, creating duplicates, or just making bad decisions overall." The result is an explosion in lines of code (LOC) that's hard for humans to keep up with, and—counter to some claims—agents can't manage the slop either.
Why LLM-as-judge doesn't work
The article dismisses the common industry approach of using AI to judge code quality. Asking a model to rate code 1–10 is "basically equivalent to a random number generator." Pairwise comparisons (A vs. B) are unstable: simply renaming the solutions can flip the model's preference. Rubrics and LLM-written tests help but are "still a far shot from actually getting rid of the slop."
The simplest metric: change in LOC
Surprisingly effective: just tracking the change in the number of lines of code. The author notes the irony: "if we started optimizing for it, it would cease to be a meaningful measure."
Verbosity
Measures duplicated and unnecessarily verbose lines. It is the fraction of lines flagged by AST-Grep or marked as clones, divided by total LOC:
Verbosity = |AST-Grep flagged lines ∪ clone lines| / LOC
Erosion
Measures how much of a codebase's mass is concentrated in a few large, complex functions. First define the mass of a function:
mass(f) = CC(f) * sqrt(SLOC(f))
Where CC(f) is cyclomatic complexity and SLOC(f) is source lines of code. Then erosion is the fraction of total mass held by functions with cyclomatic complexity greater than 10:
Erosion = ∑_{f: CC(f) > 10} mass(f) / ∑_f mass(f)These two measures—introduced by the paper SlopCodeBench—separated legacy codebases from LLM-generated slop quite well in the author's tests.
Takeaways
- LLM judges are unreliable for code quality; preference flips on renaming.
- Change in LOC is a surprisingly strong sloppiness signal, but breaks if optimized directly.
- Verbosity combines AST-Grep flags and clone detection over LOC.
- Erosion captures complexity mass in functions with CC > 10.
For teams shipping LLM-generated code at scale, these metrics offer a quantitative alternative to vibes-based evaluation.
📖 Read the full source: HN LLM Tools
👀 See Also

Nvidia reportedly developing open-source NemoClaw to compete with OpenClaw
Recent reports suggest Nvidia is working on an open-source project called NemoClaw aimed at directly competing with OpenClaw in AI development tools. The project is expected to focus on improving performance, scalability, and developer flexibility while maintaining compatibility with modern AI workflows.

Claude Sonnet 4.6 Beats Opus 4.6 on Execution in Prompt Benchmark
A Reddit user submitted a complex prompt to both Sonnet 4.6 and Opus 4.6; the Sonnet model produced a superior response judged by creativity and hidden requirements.

Local Qwen 3.6 vs Frontier Models on a Coding Primitive: Single-File HTML Canvas Driving Animation
A Reddit user pitted local Qwen 3.6 quants against frontier models (Claude, Gemini, GPT, Kimi) on a dense single-file HTML canvas driving animation task. The local Qwen 3.6-27B Q4_K_M delivered more natural motion and layering than some frontier outputs.

Benchmark Results: Qwen3.5 Models on Apple Silicon vs AMD GPUs with ROCm vs Vulkan
A developer benchmarked Qwen3.5 models (35B MoE, 27B dense, 122B MoE) across Apple Silicon Macs and AMD GPU workstations, comparing ROCm and Vulkan backends with context-scaling tests. Hardware included M5 Max, M1 Max, and three AMD GPUs with different PCIe configurations.