News & perspectives
Applied AI — min read

Why coding agents are a poor blueprint for every business task.

Software gives an AI agent tools for checking its work. Finance, operations and other business functions need their own definitions of a good outcome before similar autonomy can be useful.

Colleagues reviewing evidence in a document at a table.

A coding agent can inspect a repository, propose a change, run tests and revise the result. That sequence is compelling because it contains a feedback loop: the work is visible, some errors can be detected immediately, and a developer can inspect the proposed change before accepting it.

It is tempting to make that sequence the template for every department. Give an agent access to the right applications, tell it what to achieve and let it work. Yet software development offers conditions that much business work does not: versioned files, executable checks and people who can often examine the result in the same environment in which it was produced.

Even there, success is conditional. The original SWE-bench research uses real software issues and tests to evaluate proposed fixes, while a separate research preprint examining that benchmark found instances where weak tests allowed incomplete or incorrect patches to pass. A green test suite is useful evidence. It is not the whole definition of a good change.

Business work needs its own tests

Consider supplier onboarding. Extracting a company name and bank details from documents is a limited task. Deciding that a supplier is ready for approval requires a wider judgement: which documents are mandatory, which policy version applies, whether details agree across records and who can authorise an exception. A capable employee may know the answers from experience. An agent will need that knowledge made available through records, rules and routes to the right people.

The challenge is more than writing a better instruction. A developer can often run a test that says a function has failed. In onboarding, the equivalent test may not exist. A record can be complete but wrong. A supplier can meet the standard requirements while presenting an unusual risk. A policy can be clear on ordinary cases and silent on the one now in front of the team.

The work is to define what can be checked automatically and what needs accountable judgement. Teams can specify required fields, compare records, record the source of each conclusion and flag cases that need review. They also need examples of past exceptions, checked for whether the decision was sound and whether the policy has since changed. That creates a more useful basis for assessing an agent than a set of well-formed demonstration requests.

This is where the people doing the work become essential to engineering it. They can explain which discrepancies matter, what evidence they need before approving a supplier and when the process should stop. Engineers can turn those decisions into data flows and checks. Risk, security and commercial colleagues can set the boundaries of the process. The UK Government’s AI Playbook is written for public services, but its recommendation to begin with business and user needs rather than a technology is a sound discipline here too.

Do not borrow a productivity claim from another setting

Coding itself illustrates why leaders should resist a universal AI productivity target. A 2025 randomised study of experienced open-source developers found that the AI tools available during its trial increased task completion time in that setting. A separate set of field experiments, published in 2026, found more completed tasks among developers given a coding assistant that suggested code completions. These studies examined different tools, people and measures; neither establishes what an enterprise agent will save in supplier onboarding. METR has since reported limitations in measuring newer tools, as developers increasingly choose not to work without AI.

The implication is practical rather than sceptical. Measure the work in the setting where it will run. For onboarding, that could include time to a defensible decision, the rate of material errors, the effort spent checking documents, the handling of exceptions and any extra work shifted to legal or finance colleagues. Compare those outcomes with the existing process. A faster first draft may be valuable, but it is not the same as a faster approval.

The government’s guidance on evaluating AI interventions focuses on real-world outcomes rather than technical benchmarks and recommends a baseline for comparison. Although its remit is public-sector evaluation, that distinction helps any organisation avoid counting agent activity as business value.

Give the agent a way to fail usefully

In software, a failed test can point a developer towards a defect. A business agent needs an equally deliberate response to uncertainty. If supplier records conflict, it should show the conflict. If it lacks access to an authorised source, it should request the information or refer the case. If an exception falls outside the agreed process, it should identify the person who can decide. Completing a workflow at any cost is a poor objective.

That behaviour has to be tested with awkward cases, not merely described in a prompt. Use incomplete documents, conflicting details and policy changes. Check whether the system provides the evidence a reviewer needs and whether the reviewer has time and authority to intervene. Where it can change records or send messages, test the action itself, the audit trail and the recovery path when something goes wrong.

At Cybix, this is why the conversation starts with where intelligence belongs in the operation. A useful system may require AI, data engineering, integration and changes to the way teams work. The people who will run it need enough understanding to challenge its results and improve the process after launch.

Coding agents show the potential of capable models working with tools and feedback. The lesson for other functions is to build the conditions that make feedback meaningful. Once an organisation can say what a good outcome is, detect a bad one and act on the difference, it can decide how much of the work an agent should take on.

Contact

Make the next decision count.

No discovery funnels, no qualification calls with juniors. Write to us and a senior partner will reply, usually the same day.

Contact