Blog

2026.09.20

AI Agent Registry: Build a Control Plane in 90 Days

AI Agent Registry: Build a Control Plane in 90 Days

An AI agent registry is not merely an asset list of agent names. It is the operational record that connects business purpose, technical ownership, business sponsorship, identities, permissions, publication channels, dependent data, lifecycle decisions, and evidence. When departmental pilots multiply, the same agent is often copied into several environments, access remains after a creator transfers, and nobody can identify the production version. This vendor-neutral guide explains how buyers operating across Thailand and other sites can establish an AI agent inventory, RFP, proof of concept, 90-day rollout, change process, suspension path, retirement controls, and auditable evidence.

Executive answer: separate the agent inventory from the control plane

Use six design principles.

  1. The registry is the authoritative record of current accountability; the control plane enforces and observes runtime decisions.
  2. Separate the technical owner from the business sponsor, and ensure every production agent has an accountable business party.
  3. Record identity, permissions, channels, and data dependencies at the instance level—development, test, and production—not only at the logical-agent level.
  4. Separate discoverability from availability, deployment, pinning, and machine-to-machine invocation.
  5. Design change, suspension, and retirement when the agent is registered, including review dates, kill paths, and evidence retention.
  6. During the first 90 days, prove the discovery-to-retirement loop instead of attempting an enterprise-wide product standardization.

Current Microsoft documentation positions Agent 365 as the converged registry and control plane for discovering and managing agents, while Microsoft Entra Agent ID remains the identity and access foundation. The Microsoft 365 admin registry also groups Microsoft agents, external partner-built agents, organization-published agents, and creator-shared agents in one view. This direction is useful, but a buyer’s internal standard must not be confined to one vendor’s portal. The same vocabulary must cover other clouds, on-premises services, manufacturing APIs, RPA, MCP servers, and business SaaS.

AI Agent Registry: Build a Control Plane in 90 Days - figure 1

Why AI agent management should begin with a registry now

Traditional application inventories focus on system name, server, department, contract, and incident contact. Those fields do not explain an agent’s effective capability. An agent that only drafts text differs materially from one that sends email, writes to ERP, exercises delegated user rights, or runs autonomously overnight.

Microsoft describes an agent as an application that understands its environment or context, makes decisions, and uses available tools to pursue a goal. Core components include a model, orchestration, memory, and tools such as web search, databases, APIs, and file systems. Registration therefore cannot stop at the chat interface. It needs to connect the orchestrator, runtime identity, tools, memory, knowledge sources, channels, and downstream systems.

An invisible agent cannot be governed

Common shadow-agent patterns include:

  • an employee-built agent shared through Teams or an intranet;
  • a service identity and connector left behind after an integrator’s pilot;
  • an RPA workflow enhanced with LLM decisions but never classified as an AI system;
  • child agents called by an orchestrator but missing from the application inventory;
  • a supposedly retired job whose API key and schedule remain active; and
  • development, UAT, and production agents sharing a display name that users cannot distinguish.

Microsoft’s registry experience separately surfaces ownerless and unmanaged agents. The practical lesson is not to promise a perfect inventory on day one. Build a continuous path: discover, provisionally register, confirm accountability, connect controls, and periodically reconcile.

Identity lists and business registries are different

The Microsoft Entra list can include both agent identity objects and agents using service principals. It exposes status, Object ID, Blueprint App ID, owners and sponsors, permissions, audit logs, and sign-in logs. A business registry also needs purpose, process, audience, expected value, prohibited use, data classification, legal basis, and fallback procedures.

Do not make the identity directory the sole system of record. Link it to the registry with stable identifiers. Provisional records should include agents that lack an identity, marked identity status = missing. Conversely, an existing identity does not justify production approval if the purpose and sponsor remain unknown.

Minimum data model for an AI agent registry

Too many mandatory fields make a registry stale. Separate mandatory fields, conditional fields, and evidence links, and state who owns each field and which source updates it.

DomainMinimum fieldsGovernance question
IdentificationRegistry ID, name, instance ID, environment, versionWhich runtime instance is this?
PurposeBusiness purpose, process, audience, prohibited useWhy is it needed, and what must it not do?
AccountabilityTechnical owner, business sponsor, operator, delegateWho changes the technology and who decides continuation?
PublicationChannel, audience, discoverable/deployed/pinnedWho can find and use it?
IdentityPrincipal, authentication, delegated/autonomous mode, secret storeUnder whose authority does it operate?
AuthorizationTool, action, resource, scope, expiry, approvalWhat can it read or change?
DataSource, classification, storage, border, retention, training useWhat enters memory or output?
DependenciesModel, orchestrator, MCP/API, downstream systemWhere does a change propagate?
OperationsSLO, monitoring, correlation ID, on-call, failure procedureWho detects and handles abnormal behavior?
LifecycleRegistration, approval, next review, suspension trigger, retirementWhen and how is it reviewed or stopped?
EvidenceDesign, approval, tests, permission delta, logs, retirement proofCan a decision be reconstructed?

Separate the Registry ID from the instance ID

A business agent may have development, integration, UAT, production, and disaster-recovery instances. They share a purpose but not permissions or data. The Registry ID identifies the logical product or use case; the instance ID identifies a runtime. A procurement-query agent might be AGR-PROC-017, while its Thailand production instance is AGR-PROC-017-PRD-TH01. Cloud Object IDs, application IDs, and workload identities attach to the instance.

Keep technical ownership and business sponsorship distinct

The Microsoft Entra Agent ID administrative model distinguishes owners—technical administrators of configuration, credentials, and operation—from sponsors, who are accountable for purpose, access reviews, renewal, retention, and removal. Current documentation states that an agent identity and blueprint require at least one sponsor, while an owner is optional. Owners can be users or service principals but not groups; selected group types can serve as sponsors.

A vendor-neutral responsibility model can be expressed as follows.

RolePrimary responsibilityMust not approve alone
Business sponsorPurpose, audience, budget, continuation, residual riskCredentials and production configuration
Technical ownerConfiguration, identity, permissions, monitoring, recoveryIts own business necessity
Data ownerData use, classification, retention, transfer, deletionOverall production approval
Security/ITStandard, exception, access review, incident handlingBusiness value
Service operatorDaily monitoring, triage, evidence collectionScope expansion or high-risk change

Record delegates, succession deadlines, and an automatic suspension condition for missing accountability. If a sponsor group is used, name the representative expected to make lifecycle decisions. Microsoft notes that authorization after a dynamic-group membership change can take up to 24 hours, so emergency succession should not depend solely on group synchronization.

Separate publication channels from permissions

Teams, Outlook, Copilot, SharePoint, web, mobile, API, and batch are surfaces through which users or systems reach an agent. A shared channel does not imply the same risk when audience, tenant, country, device conditions, or operating hours differ.

  • Discoverable: visible in search or a catalog.
  • Available: eligible users may add it.
  • Deployed: administrators distribute it to an audience.
  • Pinned: it receives prominent placement.
  • Callable: an API or another agent can invoke it.

An API-only path is not “no channel”; register it as A2A/API. For a child agent, retain the caller allowlist, delegated context, rate limit, and recursion control.

What an AI control plane must enforce at runtime

An accurate registry cannot itself stop an excessive API call. The control plane consumes registry data and joins runtime identity, policy, tools, data boundaries, observability, and revocation. It may be assembled from an identity provider, API gateway, secret manager, policy engine, SIEM, and approval workflow rather than a single product.

AI Agent Registry: Build a Control Plane in 90 Days - figure 2

Unique identity and least privilege

Microsoft’s least-privilege pattern recommends a dedicated identity, named owner or sponsor, and approver for each agent, with purpose, approved data, tools, and environment documented. Scope authorization by resource, data, and action, and deny unreviewed tools and cross-tenant paths by default.

Do not write “ERP access” in the registry. Use this granularity:

FieldUseful recordAvoid
Actionpurchase_order.readERP use
ResourceThailand entity, Plant 01, approved suppliersEnterprise-wide
DataAmount, due date, item; exclude bank accountProcurement data
ModeAutonomous read; human-approved writeInherit all user rights
Duration90 days with quarterly reviewPermanent
EvidencePolicy ID, approval, test resultAgreed by email

Use short-lived tokens, just-in-time access, or action-specific approval for high privilege. Also assess aggregate effective permission: several narrow roles can combine into a broad end-to-end capability.

Allowlist tools and actions

“MCP enabled” or “plugins allowed” is too broad. Register server, tool, action, input schema, resource, maximum volume, timeout, retries, idempotency, and human approval. A search agent does not need delete. A ticket agent may receive create/update while delete and admin remain denied. Bulk update, external transfer, payment, and privilege change deserve separate gates.

Every downstream system should revalidate identity and scope; it must not trust the orchestrator alone. A sentence produced by AI that says “approved” is not authorization evidence. Connect the policy decision, API enforcement, and result with one correlation ID.

Conversation history is not a complete audit trail

Evidence needs requester, agent identity, on-behalf-of user, role, scope, tool, action, resource, policy version, approval, outcome, correlation ID, and timestamps. Retaining every prompt indefinitely creates new privacy and secret risks. Separate structured events needed for reconstruction from sensitive conversation content, with explicit purpose, masking, retention, and access.

Design four layers of suspension

One stop button is insufficient.

  1. Discovery: hide the agent from new users.
  2. Invocation: reject new sessions, API calls, and schedules.
  3. Privilege: revoke tokens, rotate keys, remove roles, and remove downstream allowlists.
  4. Data and lifecycle: handle queues, memory, evidence, outputs, legal hold, and deletion.

Distinguish a reversible incident suspension from a planned, irreversible retirement. Incident containment prioritizes immediate disablement; planned retirement prioritizes dependency migration and evidence preservation.

A 90-day roadmap for an enterprise AI agent inventory

Ninety days is a practical point to begin continuous governance, not a promise of full integration. Start with the information model and decisions, not a product purchase.

Days 0–15: discover and provisionally register

IT, security, data, procurement, and major business teams agree on a shared definition. Include systems that plan and call tools autonomously and execution units called by other agents, not only chatbots with a visible model.

Discovery sources include identity directories, cloud app registrations, API gateways, secret stores, SaaS consoles, network logs, expense and contracts, browser extensions, RPA, MCP configuration, repositories, and departmental attestations. Put automated discoveries into discovered/unverified rather than immediately treating them as approved production assets.

Deliver a registration standard, minimum fields, RACI, provisional list, impact classification, and missing-owner queue. Triage ownerless high-privilege, externally shared, confidential-data, payment, and destructive agents before less risky items.

Days 16–30: confirm accountability and risk tiers

Assign technical owners and business sponsors, then verify purpose, audience, channel, data, tools, and environment. Move ownerless items to a temporary custodian; suspend them if sponsorship is not confirmed by the deadline.

TierExampleMinimum controls
T1Public-information search and summaryTerms, citation, basic logging
T2Internal search and draftingIdentity, data boundary, owner/sponsor, access review
T3Constrained write into a business systemDedicated ID, action allowlist, approval, negative tests, stop test
T4High-impact autonomous action or privilege changeIndependent approval, JIT, enhanced monitoring, exercise, executive residual-risk acceptance

Days 31–45: connect registry and identity

Map instance IDs to cloud Object IDs, service principals, workload identities, and API clients. Create a plan for agents without identity and phase out shared credentials. Never share a production identity with development.

Automatically synchronize objective facts such as state, last sign-in, credential expiry, role, owner, and channel. Keep business purpose, prohibited use, sponsor decisions, and residual-risk acceptance under human approval. Assign a field owner so synchronization cannot erase manual evidence.

Days 46–60: map permissions, channels, and dependencies

Create a graph from each agent to tools, data sources, and downstream systems. Include user delegation, managed identities, service accounts, webhooks, batches, and queues. Start with read/write/admin, then decompose T3 and T4 permissions into actions and resources.

Confirm discoverable, available, deployed, pinned, callable, audience, and external sharing by channel. Verify that a retirement can stop APIs and schedules, not merely hide a catalog entry.

Days 61–75: prove the control plane and stop path

Select low-, medium-, and high-risk agents. Apply runtime policy from registry data. Test not only successful access but forbidden resources, expired approvals, unapproved channels, other tenants, excessive volume, duplicates, missing owners, and revoked credentials.

Use the AI agent acceptance-testing guide to structure expected, abnormal, approval, and evidence cases. Combine it with AI agent API operations for monitoring, idempotency, retries, and rate limits. This article focuses on portfolio-wide inventory and lifecycle governance rather than response-quality testing.

AI Agent Registry: Build a Control Plane in 90 Days - figure 3

Days 76–90: establish governance cadence and KPIs

Weekly, review registrations, missing owners, expired access, failed synchronization, and high-risk changes. Monthly, show sponsors necessity and use. Conduct access reviews quarterly or on a material change. Useful KPIs include:

  • percentage of production agents with unique identity;
  • percentage with valid technical owner and business sponsor;
  • percentage of permissions with expiry and approval evidence;
  • percentage of tool actions traceable end to end by correlation ID;
  • time from ownerless detection to interim containment;
  • time from disable request to invocation denial and token revocation;
  • count of expired, unused, or duplicate agents retired; and
  • percentage of material changes that completed review and retest.

RFP requirements for an AI agent registry

Specify testable capabilities and deliverables rather than product names.

1. Discovery and registration coverage

Ask how the solution discovers and deduplicates Microsoft, third-party SaaS, custom, on-premises, API-only, service-principal, identity-less, and parent/child agents. Manual CSV can assist migration, but the supplier must state the authoritative source and delta-detection method.

2. Data model and APIs

Require logical-agent and instance separation, extensible fields, history, tags, evidence links, import/export, APIs, webhooks, and change events. Contract for stable IDs, portable history, and audit retention after deletion.

3. Owner, sponsor, and approval

Require separate technical owner, business sponsor, data owner, and delegate fields; workforce-change detection; ownerless alerts; deadlines; group assignment; segregation of duties; and recertification. One generic “owner” field is inadequate.

4. Identity, authorization, and channel integration

Relate agent identities, service principals, workload identities, delegated users, credentials, roles, scopes, and tool actions. Record channel audience and discoverable/deployed/pinned/callable state, not just “Teams.”

5. Lifecycle and change management

Support states such as draft, review, approved, active, suspended, retiring, and retired, with approvers, SLAs, deadlines, and restart conditions. Detect model, prompt, tool, source, permission, and audience changes separately and trigger proportional review.

6. Audit and evidence

Preserve who changed which field from what to what, when, and who approved it. Correlate registry changes with runtime logs and export to SIEM. Verify retention, search, export, clock synchronization, masking, and legal hold.

7. Suspension, retirement, and exit

Verify session denial, schedule stop, token revocation, credential rotation, role removal, queue isolation, and webhook removal, not only catalog hiding. Include manual controls during vendor outage, migration on exit, data return, and deletion evidence.

Request a sample logical/instance schema, connector scope and permissions, demonstrations of ownerless and identity-less discovery, action-level allow/deny logs, a timed revocation test, portable exports, recovery objectives, and all licensing assumptions. Preview state and licenses change; require proof for the buyer’s tenant and contract rather than a copied feature matrix.

PoC acceptance: test the operating loop, not the list screen

The PoC should complete this entire loop:

  1. Discover and provisionally register an unmanaged agent.
  2. Assign separate technical owner and business sponsor.
  3. Connect a dedicated identity and expiring permission to production.
  4. Publish only through an approved channel and audience.
  5. Execute allowed and denied actions and trace them into downstream logs.
  6. Detect a model or tool change and return the record to review.
  7. Simulate sponsor departure and verify succession deadlines.
  8. Emergency-disable sessions, tokens, queues, and schedules.
  9. Reconcile removal of access and preserve retirement evidence.

Judge completeness, synchronization latency, false positives, false negatives, permission deltas, suspension time, audit reconstruction, and operational effort. Include duplicate names, a deleted owner, missing identity, several instances, another tenant, a shared credential, and expired approval in the test data.

Rules for change, suspension, and retirement

Review triggers

  • Prompt-only change with no capability change: owner review and regression test.
  • Model change affecting output characteristics: quality, safety, and multilingual retest.
  • Tool or action addition: permission, threat, audit, and revocation review.
  • Data source or storage change: data-owner, privacy, and transfer approval.
  • Audience, channel, or country expansion: sponsor, security, and legal review.
  • Increase in autonomy, volume, amount, or asset scope: treat as a new use case.

Suspension triggers

Missing owner or sponsor, credential exposure, unexplained privilege growth, audit loss, successful forbidden action, material data leakage, expired contract, and overdue recertification are candidates for automatic or emergency suspension. Name the decision maker and threshold at registration.

Retirement completion criteria

Stop channels, shared links, endpoints, and schedules; revoke identities, tokens, secrets, certificates, roles, and memberships; remove tools, MCP links, webhooks, queues, and downstream allowlists; handle memory, vector stores, caches, and outputs under retention policy; notify parent/child dependencies and users; retain approval and deletion proof; and confirm that the asset does not reappear in the next discovery scan.

Common implementation failures

  • Treating a one-time spreadsheet as the goal: separate synchronized facts from human judgments and process expiration continuously.
  • Making the creator the permanent owner: creators transfer and suppliers leave; define sponsor, delegate, custodian, and suspension deadline.
  • Deduplicating by display name: use stable IDs, endpoints, credentials, manifests, ownership, and tool graphs.
  • Assuming a control-plane purchase equals governance: the company still owns purpose, accountability, prohibited use, risk acceptance, and retirement decisions.
  • Assuming logs are automatically traceable: align identity, time, and correlation IDs across conversation, orchestrator, tool, and downstream API.

FAQ: AI agent management and registries

What is an AI agent registry?

It is the authoritative enterprise record of agents, their purpose, instances, technical owners, business sponsors, identities, permissions, channels, data, tools, status, and evidence. It is more detailed than an application list and carries higher-level accountability than runtime logs.

Can a CMDB or application inventory be reused?

Yes, if it can model logical agents and instances, delegated versus autonomous identity, tool actions, model/prompt/memory, channels, sponsorship, and review triggers. Connect a specialized registry where the existing model cannot represent those relationships.

What is the difference between the AI control plane and the registry?

The registry records what exists, who is accountable, and what was approved. The control plane applies identity, permission, tool, policy, monitoring, and revocation at runtime. Evaluate these roles separately even if one product offers both.

May the owner and business sponsor be the same person?

They may be combined for a small pilot, but record the responsibilities separately. Production and high-risk cases should separate the person changing technology from the person accepting continued business need and residual risk.

Are agents without identity outside the registry scope?

No. Register them as identity missing and prioritize remediation. Discovery precedes identity modernization. Production continuation should require a dedicated identity, scoped access, and a tested stop path.

How often should an agent be reviewed?

Use event triggers as well as a fixed cadence. Apply shorter expiry to higher privilege, use an organizational cycle such as quarterly for ordinary production agents, and review immediately after owner changes, tool additions, data-source changes, channel expansion, or incidents.

Can an enterprise govern all agents in 90 days?

Usually not completely. The 90-day objective is to activate definitions, discovery, accountability, a minimum schema, identity linkage, stop tests, and operating governance. Prioritize high-impact agents and keep unknowns visible for continuous improvement.

Conclusion: connect the registry to an operation that can stop

The value of an AI agent registry is not a polished list. It is the ability to answer who justifies the agent, who repairs it, which permissions and channels are really used, when it must be reviewed, and how completely it can be stopped. Separate logical agents from instances and connect technical owners, business sponsors, identities, tool actions, data, channels, dependencies, and evidence. Feed those facts into an AI control plane that enforces least privilege, allow/deny decisions, auditability, suspension, and retirement. In the first 90 days, prove one complete lifecycle on a few high-risk agents before attempting enterprise-wide platform uniformity.

TOMAS TECH can support schema design before product selection, discovery across identity/API/CMDB sources, RFP development, a 90-day PoC, operating governance, and suspension or retirement testing. You can contact us even when the organization does not yet know how many agents are active.

Primary sources

This vendor-neutral implementation guide is based on public primary sources. Product capabilities, previews, licensing, screens, and limits can change; verify current official documentation and contract terms before deployment.