AIEPCO · LLM Agent FDE☎ 18923740759
Aiepco Industrial AI Agents

A linter that silently passes is worse than none

Our static analyzer had been reporting “no issues found” on our PLC codebase for months. Then we built a fixture stuffed with deliberate bugs — and found that two of its rules had never fired on a single line of code. A post-mortem, with 8 checks you can run on your own tooling.

The short version

A checker that fails silently is more dangerous than no checker at all — it gives you false confidence. A rule written into a file is not the same as a rule that actually runs; between the two sits one thing: verifying it against a deliberately broken sample.

What a normal run looked like

[motor_control_ref.scl]
  LOC=86 SLOC=61 comment_rate=27% McCabe~9 nesting=3 magic_numbers=0
  POU: FUNCTION_BLOCK FB200_MotorControl (line 11)
  No static issues found

How we found it

We were writing a fixture — a file full of deliberately planted defects — to check whether the rules actually work. The result was unsettling: the fixture contained #m_rB used as a divisor, and unguarded array access like #m_Buffer[#i_Index]. The rules for division-by-zero risk and array-out-of-bounds risk reported nothing at all. Reading the code explained why: the pattern that looks for “a variable name after a slash” did not account for SCL local variables carrying a # prefix. Real code is #m_rA / #m_rB — after the slash comes a space, then #, then the name, and the # broke the match.

Those two rules had never matched a single line of code, since the day they were written. And the report kept saying “no static issues found”.

The siblings: five failure modes in one toolchain

Never fires
Two rules had a broken pattern and matched no code — nothing in the report looked wrong
Too broad (false positive)
VAR CONSTANT treated as a static variable: all 6 constants flagged for “non-standard naming prefix”
Noise flood
One function block produced 17 warnings of the same kind; real issues buried in the noise
Misparse
The call graph read ELSIF (...) as a function-block call, distorting call-depth analysis
Docs divorced from reality
The tooling scripts referenced in our framework documentation did not exist yet

What all five have in common: none of them raised an error. The tool did not crash, did not throw, did not warn. It quietly produced a wrong conclusion.

A fourth finding: some defects are invisible to static analysis by construction

6 / 6
Syntax / structural mutations

Missing END_IF, missing semicolon, unbalanced parentheses, naming violations, magic numbers, stripped comments — all caught

0 / 6
Semantic mutations (published on purpose)

AND → OR, deleted NOT, inverted comparison, altered initial value, removed interlock — none caught

0.9981
Text similarity after breaking it

Flipping the AND in a start condition to OR (the logic is now backwards) still scores near-perfect BLEU-4

A semantic mutation can make code look more like the reference, not less. No text- or structure-based tool can see it — so we put that limitation into the report instead of tuning the numbers.

Static analysis passing ≠ compiling ≠ being logically correct. Those are three separate sentences, and semantic correctness is the job of dynamic verification (simulation, on-machine trace comparison) and human review.

The three rules we adopted

  • Every rule needs a fixture that must trigger it: write the rule, immediately build a sample containing that defect, assert it is reported, freeze it as a regression test
  • Every rule needs a counter-fixture that must not trigger it: run it against compliant code and assert zero warnings — without this, the team learns to ignore all warnings, which equals having no checker
  • The toolchain must prove itself first: 45 self-test assertions, a few seconds to run; if that self-test is not fully green, you may not use the toolchain to judge code

What the self-test looks like

$ selftest_tools.py
assertions 45, failures 0
result: OK — toolchain behaves as declared

8 checks to run on your own tooling

  • Does every rule have a sample that must trigger it? (No → you don't know whether it runs at all)
  • Does every rule have a counter-sample that must not trigger it? (No → it may be killing valid code)
  • Is the warning count sane? (17 identical warnings in one file means noise, and the team will ignore it)
  • When the tool fails, does it error or go silent? Silent failure is the dangerous kind
  • Is there a meta-test guarding the tool itself? (Changing A and breaking B is the norm, not the exception)
  • Have you written down what the tool deliberately does not do? (Unstated boundaries are read as omniscience)
  • Does every file referenced in the documentation actually exist?
  • Before the word “passed”, is there a line saying which layer passed?

Why we publish this

There is no growth chart or ROI curve here — it is a record of our own toolchain failing and being fixed. We publish it because of a fairly plain belief: if a supplier will not say whether their own checkers have been verified, you should trust what they hand you even less. Saying three separate sentences — static checks passed, it compiles, the logic is correct — is more useful to the person paying for the work than one vague “quality assured”.

Put your production programs on a verifiable engineering process