AIEPCO · LLM Agent FDE☎ 18923740759
Aiepco Industrial AI Agents

We publish the evidence — and the boundaries

Everyone claims accuracy. We would rather state clearly which layer was checked, which was not, and when our tools last checked themselves.

Evidence 1: our own tooling once failed silently

Field Note 001|Two rules in our verification toolchain had never fired

A static checker that had run for a long time kept reporting “no static issues found”. Only after we built a deliberately defective sample did we discover that the divide-by-zero and array-bounds rules had never matched a single line since they were written — the regex ignored the # prefix on SCL local variables.

That led to three more of the same disease: rules too broad (constants misjudged), noise flooding (17 identical warnings on one block), and a parser error (treating ELSIF (...) as a block call). What all five failure modes share: none of them raise an error.

Full text: WeChat account 「智能体驻场日记」

Evidence 2: mutation testing — what our tools cannot catch

6 / 6
syntax / structural defects

missing END_IF, unbalanced parentheses, missing semicolon, stripped variable prefixes, magic numbers, low comment ratio — all detected statically (80% threshold)

0 / 6
semantic defects (published on purpose)

AND to OR, removed stop condition, inverted comparison, changed constant initial value, removed fault latch, removed NOT — not one detected statically

The clearest example: changing AND to OR in the start condition (reversing the logic) still leaves text similarity at BLEU-4 0.9981. We write that conclusion into our reports instead of making the numbers look good: static analysis passing ≠ compiling ≠ logic being correct. Semantic correctness belongs to dynamic verification (simulation / real-machine trace comparison) and human review.

Evidence 3: before judging your code, prove the ruler is accurate

  • 45 tool self-test assertions covering: every rule must have a defect sample that triggers it, the reference implementation must produce zero warnings, mutation detection must meet the threshold, and every script referenced in the docs must actually exist
  • The rule: if the tool's self-test is not fully green, it may not be used to judge code
  • When a tool is unverified, its “pass” is not evidence — that is exactly where the failure story above came from

Example scenarios (illustrative — not client projects)

Example A|Water treatment dosing
Dosing rate from flow and pH feedback, with bumpless manual/auto transfer, reagent-low interlock, pump fault latching and graded alarms.
Example B|Cleanroom AHU interlocking
Supply fan start/stop sequence, differential pressure interlocks, filter blockage alarms and maintenance prompts; key data reported to the host system.
Example C|Packaging line cycle-time retrofit
New station integration, recomputed cycle time, upstream/downstream interlock refactoring, with open-items log and on-site confirmation checklist.

The scenarios above are illustrative — they show how we would approach such work — and are not client projects. Real anonymized cases will replace them once projects complete and publication is authorized.

Put your production programs on a verifiable engineering process