How Capital One evaluates, governs its AI agents
Capital One walked through how it is evaluating and governing AI agents in production at scale. The company built its own system architecture to wrangle its agents.
Speaking at CoreWeave Fully Connected 2026, Maulin Patel, Managing VP of Product, AI and ML platforms at Capital One, walked through the company's agent evaluation processes in a regulated environment.
Capital One is a bank that sees its technology stack as its differentiator. The company is using its technology platform to integrate acquisitions, most recently the purchase of discover, create efficiencies and drive growth. Capital One CEO Richard Fairbank said at a recent investor conference.
"We have focused on entering and building a lot of scale in inherently attractive businesses with good growth and earnings power prospects. Along the way, we've built a business model, leveraging technology and analytics and measurement, to be able to have high octane businesses."
- CoreWeave launches CoreWeave Forge: Here's a look at the strategy
- Caterpillar's AI autonomy efforts accelerate, but domain knowledge drives returns
Fairbanks said Capital One has been investing heavily in technology for 14 years. "We have been working backwards from the dramatic transformation of the business marketplace with modern technology, data, and AI," said Fairbank on the company's second quarter earnings call. "We continue to invest in some very powerful foundational capabilities as well as AI infrastructure and specific AI experiences."
That backdrop is notable given that Capital One has a tendency to push the edge. Patel noted that AI agents are scaling across the business and that means a lot of evaluations. Agentic AI is an evolution of Capital One's cloud and data transformation. Patel also noted that AI agents serving 100 million customers requires processes that deliver AI safely in a regulated environment.
"We serve more than 100 million customers. At that scale we cannot govern autonomous AI agents just by raising chats," said Patel. "We built an evaluation first architecture designed to deliver trust-moving outputs."
Key ingredients of Capital One's AI agent evaluation stack include:
- Scoring metrics and using moving averages to capture statistical depth and reasoning through workflows.
- Capital One uses an LLM to judge the chain of AI agents.
- Contextual accuracy is measured to catch hallucinations.
- The models that judge AI agents are supplemented and calibrated with humans.
Patel said to pull off AI agents evaluation at scale, Capital One said code and CoreWeave collaboration is essential. He also added that human in the loop is also critical. "Human oversight is super important, especially in regulatory industries," said Patel.
"Our subject matter experts audit complex financial logic and calibrating our automated LLM channels. This architecture works really well for us for single LLMs and communications. However, autonomous mobile agents introduce entirely new set of players," Patel added.
Among the issues Capital One has seen:
- Capital One would find that AI agent conversations would lose context and forget the core intent over multiple terms.
- Models would also contradict decisions made earlier.
- The other issue is circular logic and the AI agent immediately asking for the same information.
- And finally, early miscues in multi-step tasks can compound.
"Those challenges led us to build our evaluation architecture from the ground up," said Patel.
Capital One's evaluation architecture is also designed to enable developers and not hinder them. "Embedded governance is an accelerator and when you embed continuous evaluation directly into the developer library, governance becomes an accelerator because your developers can scale safely," said Patel.