This time final yr, I launched a Deeplearning.ai course with Andrew Ng on Governing AI brokers with the aim of teaching builders on the fundamentals of knowledge safety for efficient and accountable agent deployment. The motivation for the course got here from IBM’s 2025 breach research, which discovered 97% of the organizations that suffered an AI-related breach lacked correct AI entry controls, and 63% had no AI governance coverage in any respect.
A yr is an eternity in AI improvement, and we have now come a great distance from brokers with zero governance or grappling with governing a single agent. Groups at the moment are confronting a brand new governance problem: agent sprawl (the uncontrolled development of autonomous AI brokers throughout a corporation with out centralized monitoring, possession, or governance). In response to Gartner, by 2028 the common Fortune 500 enterprise will use over 150,000 AI brokers. Nonetheless, in response to the agency, solely 13% of organizations imagine that they’ve the suitable AI agent governance in place. Since final yr, the flexibility to deploy has gotten simpler than ever. The onset of coding brokers like claude code and codex, in addition to a wide range of low-code/no-code choices have lowered the barrier to agent deployment considerably. One result’s groups at the moment are confronted with the potential for their agent fleet participating in massively wasteful token utilization and incurring unexpected prices. One other severe problem is the extra pathways for delicate knowledge to leak out. Unsurprisingly, governance challenges have advanced.
Presently, I am watching clients construct brokers sooner than ever, and the form of the issue has shifted. A yr in the past we targeted on including the 4 pillars of governance to a single agent: lifecycle administration, danger administration, safety and observability.
We knew even then that constructing an agent wasn’t the exhausting half. You may arise an agent in a day, wire in observability, level it at a copied-over slice of knowledge, and it seems to be as if it’s production-ready. You then attempt to run it for actual, towards stay methods and at scale, and also you hit the wall that truly issues: infrastructure at scale. Now, this hasn’t modified, however what’s the shift? The distinction is that is occurring with dozens of brokers/sub-agents on the similar time. Brokers are multiplying sooner than quite a lot of groups are capable of govern them. Groups now must be outfitted with an information platform that scales centralized monitoring, permissions, and governance with the variety of brokers being constructed. Let’s study how the pillars of governance have advanced in 2026.
The place we had been: 4 pillars and one agent
Final yr I taught this free course by way of DeepLearning.AI on governing AI brokers, the place we constructed a ruled HR analytics agent on Unity Catalog and MLflow. The entire course distilled into 4 pillars:
-
Lifecycle Administration (Separation of Duties): model, deploy, and retire brokers with full lineage throughout dev, staging, and prod.
-
Danger Administration (Protection in Depth): overlapping defenses equivalent to PII detection, guardrails, compliance controls, and monitoring, from knowledge ingestion by way of mannequin efficiency.
-
Safety (Least Privilege Entry): brokers and customers get solely the minimal permissions their function requires, enforced by way of authentication, encryption, and granular entry controls.
-
Observability (Audit Every thing): log each enter, output, and choice for full traceability and compliance.

The pillars sound summary till you sit within the room with authorized, auditors and management. Then they collapse into three very concrete questions.
-
What can the agent attain? In our construct, the reply was: no direct desk entry, ever. Every thing ran by way of layers, from masking to views to teams to features, with knowledge classification enforced at each and aggregation-only entry on delicate tables. In apply meaning the agent can reply “What’s attrition in engineering this quarter?” whereas being structurally incapable of surfacing anybody individual’s wage.
-
What modified, and may I undo it? The agent was registered as a versioned Unity Catalog mannequin and deployed from that model, sitting on high of version-controlled features and views. So when reply high quality drops on a Tuesday, you already know precisely what shipped on Monday, and you’ll revert it as a substitute of debugging a black field in manufacturing.
-
Can I reconstruct what occurred? Each perform name was logged, MLflow traced each run, and there was an audit path from the question all the best way all the way down to the underlying knowledge. When an auditor asks what the agent touched on March third, that is a question, not a three-week investigation.
However discover the scope of all this. The complete recreation, a yr in the past, was getting one agent safely into manufacturing.
The place we at the moment are: governance has to scale
Whereas the 4 pillars nonetheless maintain, now each considered one of them has to use throughout a complete fleet of brokers directly. That takes two issues: insurance policies that apply globally, to each agent, and the infrastructure to implement it.
Each agent wants ruled entry to stay knowledge, not a copied-over pattern that seemed advantageous within the prototype. Additionally, each agent generates its personal report: traces, spending, and entry logs; and all of that has to land someplace you’ll be able to question. Now, hand-configuring lineage and least-privilege for one carefully-built agent merely would not survive contact with 100 of them working on an information platform that was by no means wired for it. The exhausting work strikes down a layer, from the agent to the infrastructure it runs on.
The excellent news is that the infrastructure to assist governing brokers at scale is changing into a high precedence of firms which can be constructing brokers and wish to mitigate danger– in addition to a precedence for firms who’ve beforehand skilled unauthorized acts and breaches of obligation from their brokers. The reply to sprawl is a management airplane: one ruled layer that each agent runs by way of, sitting on high of the identical knowledge platform that already holds your tables, permissions, and lineage. On Databricks, that layer is Unity Gateway, a system that governs how builders attain AI coding brokers, fashions and instruments. Admins configure and govern centrally; builders simply run a command, ug claude or ug codex, and get an accredited agent with the suitable settings already baked in. 4 capabilities do the heavy lifting.
-
Agent Configuration. Admins outline the approved set, which fashions, MCP servers, abilities, and budgets a group can use, then publish it. Builders set up the Unity Gateway CLI as soon as and launch accredited instruments with ug claude or ug codex; each launch checks for native drift and enforces the printed guidelines. Governance stops being one thing you wire into every agent and turns into an org-wide default builders inherit routinely.
-
Good Routing. As an alternative of sending each request to the largest accessible mannequin, the gateway matches job complexity to mannequin functionality: low-cost fashions for easy work, highly effective ones for the exhausting issues. On Databricks’ personal inside coding benchmark, good routing alone produced a 35% value saving. Governance now quietly decides which mannequin runs on which job.
-
Good Budgets. Spend turns into first-class. You set month-to-month budgets, shared throughout a group or per person, and resolve what occurs at every threshold: ship an alert, block additional requests, or each. The gateway may nudge towards cheaper choices as spend climbs, recommending a smaller mannequin or a low-cost open mannequin when you cross, say, 80% of finances. Price, which used to floor in a finance spreadsheet weeks later, turns into one thing you govern in close to actual time.
-
Unified Tracing. Each device name is traced routinely: its identify, arguments, errors, token counts, and latency land in a unified desk you’ll be able to question. That turns value management right into a lookup as a substitute of an investigation. In one case, Databricks traced roughly $499,000 a yr in wasted tokens to seven small bugs in device servers, and stuck them in about an hour.
Governance scales the identical manner, by way of coverage. Insurance policies use attribute-based entry management (ABAC): you write one rule towards ruled tags and it applies all over the place, so an MCP server may be permissioned all the way down to particular person instruments, and entry keys off the attributes of the person or agent making the decision. Service insurance policies add guardrails on each request and response, blocking or masking delicate knowledge and catching immediate injection, unsafe content material, and hallucinations earlier than they attain a person. A brand new agent inherits the suitable permissions and guardrails from who it belongs to, not from a bespoke grant.
This is not a principle. One buyer, Concurrence, described routing all their visitors “by way of a single ruled path whereas sustaining identity-level attribution and entry to accredited fashions and MCP instruments,” a deployment that ran 61 billion enter tokens throughout roughly 360,000 requests. That is the distinction between governing an agent and governing a corporation’s complete agent footprint.
The shift: similar pillars, greater floor
The 4 pillars did not get changed, they scaled to cowl a fleet and we added a fifth pillar: value.
|
Pillar |
Then: one HR agent |
Now: a fleet of coding brokers |
|
Lifecycle Administration |
Model and deploy one agent through MLflow + UC |
Central agent configuration printed to the fleet, with permissions and lineage on each mannequin, MCP, and abilities |
|
Safety |
Least-privilege and masking on one dataset |
GRANT/DENY, ABAC, and contextual insurance policies present dynamic entry controls to fashions, providers, and tooling. Granular MCP device prevents unapproved device utilization with out eradicating all MCP performance. |
|
Observability |
MLflow traces and classes |
Unified hint tables and dashboards throughout the entire fleet |
|
Danger Administration |
Catch failure modes earlier than manufacturing |
Service-policy guardrails on each request and response: delicate knowledge, immediate injection, unsafe content material, hallucinations |
|
Price |
Over-engineered brokers and lack of value insights led to shock billing, usually stifling improvement. |
Good routing for mannequin effectivity, value visibility, fee limits and finances caps to forestall tokenmaxxing whereas selling valuemaxxing. |

In apply, the pillars now map to concrete capabilities delivered by way of Unity Gateway, Unity Catalog, and MLflow: central agent configuration, insurance policies (guardrails and ABAC grants), unified traces, and good routing and budgets. The one-line model: governance went from gatekeeper to regulate airplane. A yr in the past the query was “can this agent see this row?” Now it is “which mannequin runs, on which job, at what value, underneath whose identification, throughout each agent within the firm?”
What’s subsequent
The course closed on a roadmap: blue-green deployments for brokers, a centralized gateway for endpoint administration, and anomaly detection on agent habits. It is satisfying to observe these transfer from “coming quickly” to “transport.” The gateway layer, specifically, is now actual infrastructure fairly than a slide.
In case you’re constructing brokers, the basics have not modified. The 4 pillars are nonetheless the on-ramp, and the free 75-minute course nonetheless walks you thru them finish to finish. What’s modified is the ceiling. Begin by auditing your personal brokers towards the pillars. Then ask the larger query: not simply whether or not every agent is ruled, however whether or not you’ll be able to steer all of them directly. As a result of in case your agent deployment is caught ready on a compliance evaluation, the blocker most likely is not the mannequin.















