← All resources

GUIDE

Do You Really Know What Your Teams Are Doing With AI?

Companies are asking everyone to use AI while struggling to see what their teams have built. Visibility reveals what deserves more investment, what should become a shared standard, and where accountability is missing.

A company can know exactly what it spends on AI and still have very little idea what it is buying.

It can count licenses, monitor token consumption, publish an acceptable-use policy, and ask every department to demonstrate progress. Those are all useful activities. They can also coexist with a remarkably incomplete understanding of the work itself: what people have built, which processes now depend on it, whose approach is worth replicating, and what happens when something changes. An organization can become quite sophisticated at administering access to AI while remaining almost incurious about the capabilities taking shape behind that access.

We keep coming back to a simpler question: Do you really know what your teams are doing with AI?

The uncomfortable word is “really.” Most leaders can point to a handful of projects, an enthusiastic team, or a compelling demonstration. The harder task is explaining what exists beyond those examples, what is actually being used, and which of those experiments have quietly become part of how the business operates. The distance between those two answers is where a substantial management problem is developing.

In conversations we have been having, a familiar sequence keeps emerging. AI begins as something employees need permission to try. Security has questions. Legal has questions. A small group gets access. Then the mandate changes: move faster, use AI, find efficiencies, demonstrate adoption. The ambition accelerates much faster than the organization’s ability to understand what that ambition is producing. People respond exactly as they have been asked to respond. They experiment, improvise, share what works, and build around whatever tools are available.

That initiative is valuable. It also produces an accumulation of business logic that the company may struggle to find, let alone manage.

Consider an ordinary customer support workflow. Someone develops instructions that help an assistant classify an incoming issue, consult the right documentation, and draft a response. A colleague copies them and adds an exception for a particular customer segment. Another team adapts the workflow for a different product. Eventually, one version can update a record or initiate an escalation. Each change is comprehensible in isolation. Taken together, they create several related capabilities with different assumptions, permissions, and consequences.

The company may still describe all of this as “using an approved AI tool.” That description tells us almost nothing about the process it now relies on.

Approving a platform answers a question about the platform. Understanding what employees build within it requires another level of attention. Prompts, skills, instructions, agents, workflows, tools, and integrations encode decisions about how work should happen. They can contain the accumulated judgment of someone who understands a difficult process exceptionally well. They can also preserve an outdated assumption with extraordinary fidelity.

We use the term capability for these reusable ways of getting work done with AI. The terminology is secondary. A leader should be able to ask who owns the customer escalation workflow without first learning which vendor calls it a skill, an agent, or a project. The business needs a coherent account of its work even when its technology suppliers insist on different vocabularies.

Once that work becomes visible, much more useful questions become possible. What have we already built? Who is responsible for maintaining it? Which version has been reviewed, and where is it available? Are people using that version or a derivative that changed three weeks ago? Are several teams independently solving the same problem? When a capability appears successful, what evidence supports that conclusion?

Those questions reveal opportunity as readily as they reveal risk. A particularly effective workflow might be saving one team hours while another team struggles through the same task manually. Someone may have developed an excellent method for checking a document, preparing an analysis, or investigating an exception, but its distribution still depends on knowing the right person to ask. The company has paid for the learning. It has yet to make that learning available to everyone who could benefit.

There is a peculiar inefficiency in asking an entire organization to innovate while leaving each team to rediscover the useful parts.

Duplication deserves more judgment than a simple count, though. Two teams approaching the same problem differently may be conducting precisely the experimentation the company needs. Their circumstances may differ. One approach may expose weaknesses in the other. Visibility gives the organization a chance to distinguish productive variation from redundant effort, examine the results, and decide what deserves to become a shared standard. Without it, both good experimentation and waste remain equally obscure.

The same problem appears in how we evaluate AI investment. Token consumption is easy to count, which makes it tempting to treat as a proxy for progress or excess. But a large bill could represent valuable work, repeated failures, unnecessary complexity, or some combination of all three. A small bill could indicate efficiency or an initiative that nobody uses. Neither number explains itself.

Suppose a team spends more on a workflow that materially reduces the time required to resolve a customer issue while maintaining quality. Cutting its budget indiscriminately could constrain one of the better investments in the business. Another workflow might consume fewer resources and still be a poor investment because its output requires extensive correction. The sensible allocation depends on the work being performed and the result being achieved.

We should be equally suspicious of attempts to turn that complexity into an unexplained productivity score. A frequently retrieved capability is attracting interest; retrieval alone does not establish successful execution. Repeated execution shows activity; it does not, by itself, demonstrate accuracy, time saved, or commercial value. Those conclusions require evidence from the process and the people responsible for it. A useful view of AI adoption makes these differences explicit and leaves uncertainty visible.

The objective is to give leaders enough context to exercise judgment: where to invest, what to improve, which work to share, and which assumptions deserve scrutiny. The people producing the most valuable capabilities may need more support, more distribution, or more room to experiment. Finding them is an offensive advantage as much as an administrative responsibility.

The accountability questions become harder as those capabilities acquire greater authority. Drafting an internal summary and changing a customer’s account are different acts, even if both happen through the same model. The relevant context includes the instructions in effect, the information available, the tools accessible, the permissions granted, and the point at which someone was expected to review the result. Merely naming the model leaves most of that story untold.

Imagine trying to investigate a consequential mistake without knowing which version of a workflow was used. The instructions available today may differ from those in effect at the time. The original author may have moved teams. A colleague may have altered an exception, expanded access, or removed a review step while fixing an unrelated problem. Reconstructing the sequence from messages and recollections becomes an exercise in institutional archaeology, conducted at exactly the moment the business needs a dependable account.

Provenance gives that investigation a starting point. Ownership, version history, approval records, and deployment context help establish what the organization intended and what it made available. Execution records and evaluations, where captured, help explain what actually occurred. None of those records makes a probabilistic system deterministic, and an approval does not guarantee a correct outcome. They do give the people responsible something more substantial than a reassuring policy and the current contents of a folder.

That is why visibility has to come first in the operating model for AI governance. A policy can establish boundaries before an inventory is complete. But applying those boundaries intelligently requires knowledge of the work they are supposed to govern. A requirement for human review has little operational meaning if nobody can identify the workflows that need it, the versions that include it, or the changes that might have undermined it.

Visibility itself has to be credible. Connecting a repository or an AI platform does not establish that every relevant capability has been discovered. An inventory should make its coverage legible: which sources are connected, when information was last refreshed, where activity can be observed, and where the organization still lacks evidence. Absence from a dashboard is not evidence of absence from the business. Concealing that limitation would reproduce the original problem with better graphics.

Employees also need a reason to participate. If making their work visible creates only additional scrutiny, they will have little incentive to maintain an accurate record. The system should help them find useful work, establish authorship, get an appropriate review, distribute improvements, and avoid rebuilding something a colleague has already solved. Accountability is much easier to sustain when it accompanies practical assistance.

From there, standardization becomes a deliberate progression. A useful experiment gets an owner and a clear purpose. Its assumptions and limitations become explicit. A particular version is reviewed for an intended use, made available to the right people, and maintained as those people learn from it. Necessary variations remain understandable because their relationship to the shared version is preserved. Changes can be examined, and drift between an approved version and an external implementation can be addressed.

The organization begins to retain the benefit of experimentation. A person’s discovery becomes something a team can depend on, improve, and pass on without losing its history along the way.

This is the problem we are building Commonset around. Commonset gives teams one place to see what AI capabilities exist, who owns them, and which ones the organization actually trusts and uses. Ownership, provenance, review, versioning, access, distribution, and drift belong in that account because they help people make decisions about the work. The technical description is a control plane for reusable AI capabilities. The business purpose is to make what the organization is building understandable and manageable across teams and platforms.

There is still a great deal to learn about which forms of AI use produce durable value. Companies need room to discover that through actual work. They also need a way to retain what they learn, recognize where it applies, and examine the consequences as experiments become dependencies. Otherwise, every new wave of adoption adds activity without necessarily adding institutional understanding.

The next useful conversation about AI inside a company should be specific enough to change a decision. Show us a capability another team should be using. Show us an important workflow whose ownership is ambiguous. Show us where versions have diverged, where investment is paying off, and where the evidence is still insufficient. Then decide what to expand, what to standardize, and what to repair.

A company that can answer those questions can put more resources behind its best work and take responsibility for the processes emerging around it. That is a far more consequential achievement than being able to say everyone has an AI account. It means the organization is beginning to understand what it has built.