Sterling Financial: AI Attributions on COBOL in 2026

Listen to this article · 11 min listen

Anya Sharma had a classic 2026 problem. As CTO of Sterling Financial, a regional bank running on COBOL systems from before her parents were born, she had to plug in a new suite of AI financial advisors. These things were supposed to offer personalized investment strategies, but her mandate was firm: don’t touch the core infrastructure. This created the immediate, show-stopping problem of AI agent attribution. If one of these new AI agents changed a customer record, how could anyone prove it inside a system that was never designed to know what an AI even was? The audit trail would be a black hole.

Key Takeaways

  • You can’t get AI agent attribution in a legacy environment without a dedicated middleware layer. It’s the only way to translate modern API calls into old-school protocols like MQSeries or SOAP.
  • Logging and auditing for AI actions isn’t optional. You have to capture the agent’s ID, a timestamp, and exactly what data it changed, and you have to do it within your existing database structures.
  • Run daily batch reconciliation. These data integrity checks are your safety net for validating that what the AI *thought* it did actually happened correctly in the source system.
  • Your security model has to include the AI agents themselves. Use standard auth like OAuth 2.0 or SAML to validate their identity before they ever get near a legacy system.
  • Go slow. A phased integration with regression testing at every single step is the only way to avoid breaking something in a stable, production-critical application.

The Challenge of Digital Ghosts

Plenty of companies are in Sterling Financial’s shoes. If you’re in finance, healthcare, or government, you’re likely sitting on legacy systems that are absolute workhorses, secure, fast, but built for a world before distributed AI. They talk in forgotten languages like IBM MQSeries, SOAP, or even flat file drops. For Anya, the issue wasn’t just making the new AIs and the old systems communicate. The real nightmare was attribution. Every single action an AI agent takes has to be traceable back to that agent inside the legacy data itself, because if you can’t produce a perfect audit trail for the regulators, you’re facing massive fines.

“Our COBOL systems are like highly efficient, silent workhorses,” Anya told her team. “They do their job perfectly, but they don’t speak ‘AI.’ We can’t just plug in a new API and expect it to understand a flat file or a CICS transaction.” The bank’s audit logs were great for tracking human actions, but they had no place to record an AI agent’s ID, its confidence score, or the ‘why’ behind its decision. That gaping hole meant they couldn’t comply with the Federal Reserve’s new guidance on AI in banking from 2023. It was a non-starter.

Building the Translation Layer: Middleware as the Rosetta Stone

A full-scale migration was out of the question. David Chen, the senior architect on Anya’s team, knew the cost and risk of replacing the core COBOL systems were astronomical. You don’t just rip out banking systems that have been running reliably for decades. The only realistic path was a clever middleware solution. This new layer would be the project’s Rosetta Stone, letting the modern AI agents (built on platforms like DataRobot and H2O.ai) talk to the COBOL apps without anyone having to touch the core legacy code. David’s team built the middleware to show the AI agents a clean set of RESTful APIs. So, when an agent wanted to update a customer’s profile, it just made a standard API call.

That simple API call triggered a complex chain of events inside the middleware. It first checked the agent’s identity with an OAuth 2.0 token, no auth, no entry. Then came the real magic: it translated the JSON payload from the API call into whatever the COBOL program needed, which could mean building a fixed-length record or a specific CICS transaction message from scratch. The final, and most important, step for attribution was injecting metadata directly into the data going to the legacy system. The middleware would stuff in a unique AI agent ID, a timestamp, the AI model version, and a reference ID for the decision log. This meant hunting for unused fields in the old COBOL copybooks or, when they had to, adding small auxiliary tables to hold the new attribution data.

The Attribution Imperative: Logging and Auditing for Compliance

Getting the data across was one thing, but proving it for compliance was another. For AI agent attribution to mean anything, you can’t just log that “an AI” did something. Regulators want to know which AI, at what time, based on what data, and exactly what it changed. To satisfy this, Sterling’s team set up two parallel logging streams. First, the middleware itself logged every single interaction it handled, the full request, the response, and any errors during translation, into a modern database that the ops team could search in real time.

At the same time, they had to get that attribution data into the legacy system itself. “We couldn’t rewrite half a million lines of COBOL overnight,” David admitted. “So, we went treasure hunting.” His team found existing filler fields and unused flag bytes in the old record layouts and repurposed them to hold the AI agent IDs and transaction references. It was a hack, but a careful one. For the most critical changes, they performed surgical modifications, adding small logging routines directly into the COBOL programs to write to a dedicated audit file. This two-pronged approach, making tiny, precise changes to stable code, was the only way to avoid introducing new bugs into the bank’s core systems. All of it, the middleware logs and the COBOL audit files, got ingested into a central Amazon Redshift data warehouse, which let auditors connect the AI’s decision log with the final transaction record on the mainframe.

Data Integrity and Reconciliation: Trusting the AI’s Handiwork

Of course, this whole setup introduces a new kind of risk: data drift. What happens if the AI agent fires off an update, but the legacy system chokes on it or the middleware translation has a bug? Anya knew that data integrity was everything, so her team built a safety net: daily batch reconciliation jobs. Every night, a process would run that compared the actual data in the legacy system against the intended state recorded in the middleware logs. If anything didn’t match, it fired an immediate alert for the operations team to investigate manually.

Take a portfolio rebalance, for example. The AI agent would recommend the new allocation, and the middleware would log it. The next morning, a reconciliation script would automatically query the COBOL investment system to confirm the trades were actually executed and the client’s portfolio now matched the AI’s plan. Anya was clear about the purpose. “We have to guarantee the entire chain of custody, from the AI’s decision all the way to the final record in the legacy system, is perfect and provable,” she said. “We’re talking about people’s money, so there’s zero room for error.” These daily checks were lifesavers, catching small integration bugs before they could blow up into massive data corruption issues.

Working through Security and Scalability: A Continuous Evolution

Authentication was just the first step for security. The Sterling team locked down the AI agents with strict access controls based on the principle of least privilege, meaning each agent could only perform the exact transactions it was built for and nothing more. Every hop of the communication, from agent to middleware to legacy host, was encrypted with TLS 1.3. They also brought in outside pen testers to hammer on these specific integration points, actively trying to compromise an agent’s identity or sneak bad data past the middleware.

The other big concern was scalability. The middleware couldn’t become a bottleneck as the bank rolled out more and more AI agents. David’s team planned for this from day one, building the middleware as a set of microservices that could be scaled independently. They ran the whole thing on Kubernetes, which gave them the power to spin up more resources automatically when traffic spiked. It was a lot of upfront work, but it meant they built a foundation that could support years of future AI projects instead of just a one-off integration.

The Resolution and Lessons Learned

The project worked. By the end of 2026, Sterling Financial had its new AI advisors running, fully integrated with the old COBOL systems. Thanks to the careful work on the middleware, logging, and reconciliation, every action taken by an AI was completely attributable and auditable. Anya’s team managed to bring modern AI capabilities to the bank while protecting the stability and security of its most critical operations.

The big takeaway from Sterling’s experience is that integrating AI with legacy systems is a risk management and compliance project first, and a technology project second. You need people who deeply understand both modern AI and how a 40-year-old piece of software actually works. That intelligent translation layer, the bulletproof audit trail, the constant data validation, those aren’t nice-to-haves. They are the absolute price of admission for using AI without risking your core business or getting slammed by regulators.

If you’re looking at a similar project, the advice is straightforward. Spend the money on the middleware. Obsess over granular attribution. Validate everything, twice. Your stable legacy systems are a huge asset, and making sure your AI can play nicely with them is how you keep them that way.

What are the primary challenges of integrating AI agents with legacy systems?

You’re fighting battles on a few fronts. The protocols don’t match between new AI and old systems. You have to create a bulletproof audit trail for every AI action to satisfy regulators. You need to constantly check that the data remains consistent across both systems. And you have to secure the entire chain, from the AI agent all the way to the mainframe.

How can organizations achieve AI agent attribution in systems not designed for it?

The best approach is to build a middleware layer specifically for this. This layer intercepts the AI’s action and injects attribution data, like an agent ID, timestamp, and model version, into the transaction before it hits the legacy system. You can often repurpose old, unused fields in your legacy data structures to hold this information, making every action traceable.

What role does middleware play in legacy system integration for AI?

It’s a translator and a gatekeeper. The middleware converts modern API calls from an AI agent into the old protocols and data formats that legacy systems like COBOL or CICS expect. It also handles the security (authentication and authorization) and injects the metadata needed for attribution. Essentially, it lets the AI agent live in a modern API world, ignorant of the legacy complexity it’s talking to.

What security considerations are critical when AI agents interact with legacy systems?

First, your AI agents need their own identities and strong authentication, like OAuth 2.0. Second, lock them down with the principle of least privilege so they can only do their specific job. All communication must be encrypted (e.g., TLS 1.3). Finally, you need to hire people to regularly attack the integration points through penetration testing to find holes before someone else does.

Why is data reconciliation important for AI agent integration with legacy systems?

It’s your safety net. Data reconciliation proves that the changes an AI intended to make were actually processed correctly by the old system. By running regular checks that compare the AI’s “sent” log with the legacy system’s “received” state, you can catch errors, data drift, and other inconsistencies before they become a huge mess. It’s fundamental for data integrity and for proving your system works as advertised.

John Thornton

Principal AI Ethics and Attribution Scientist Ph.D. Computer Science, Carnegie Mellon University; Certified AI Ethics Professional (CAIEP)

John Thornton is a leading AI Ethics and Attribution Scientist with 15 years of experience specializing in the provenance and accountability of autonomous agents. Currently a Principal Researcher at Veridian Dynamics, he spearheads initiatives to develop robust frameworks for identifying the origin and intent of content. His groundbreaking work on the 'Thornton-Veridian Attribution Model' is widely cited for its innovative approach to tracing complex AI decision-making chains. He is a frequent speaker at industry conferences and a published author on the ethical implications of advanced AI systems