ARTICLE
AI Skills and Enterprise Governance: Statistics and Research
What the evidence says about AI skill performance, versioning and review, with broader workplace AI data clearly separated from skill-specific findings.
Reviewed: 2026-10-07
A useful AI skill needs more than a memorable name. Before a team relies on it, someone should be able to explain what it helps with, where it came from, how it was evaluated, and which version the team should use.
This research hub separates evidence about AI skills from broader evidence about workplace AI. Benchmarks measure performance under particular test conditions. Platform documentation describes supported behavior. Employee and executive surveys report experiences and perceptions. Those are different kinds of evidence, and none independently demonstrates demand for a particular governance product.
Commonset curated this collection. We did not conduct the cited studies. The practical recommendations below are our interpretation, not findings attributed to the researchers.
Evidence about AI skills
1. Curated skills improved average benchmark performance, but not every task
SkillsBench v4 reports an average pass rate of 33.9% without skills and 50.5% with curated skills, a 16.6 percentage-point increase, across 87 tasks and 18 model–harness configurations. Thirteen tasks showed negative skill deltas.
Source and limits: SkillsBench, v4, June 14, 2026. This is a containerized agent benchmark using curated material, not a workplace productivity or return-on-investment study. Results do not establish the performance of arbitrary skills.
Commonset interpretation: Evaluate a candidate skill against the task it is supposed to improve, including a baseline without it. Keep the tested revision with the results. A strong average result is not a reason to make every skill available for every task, and a successful demonstration should not substitute for checking failure cases.
2. The skill format does not establish organizational approval
The Agent Skills specification requires name and description in the SKILL.md frontmatter. Its optional metadata field can hold information such as author and version. The specification does not require an organizational approval record.
Source and limits: Agent Skills specification, checked October 7, 2026. This is a living technical specification, not a survey of what organizations actually record. Providers and organizations can add controls beyond the format.
Commonset interpretation: Treat format validation and approval as separate decisions. For a shared skill, record who is responsible for it, the exact revision reviewed, its intended audience, and any conditions on its use. Do not infer those decisions from a filename or the fact that an agent can load the package.
3. A skill review needs to cover more than the instruction file
Anthropic documents that skills can include executable code. Its security guidance recommends trusted sources and, for less-trusted skills, auditing bundled files, dependencies, resources, and instructions or code that contact external network sources.
Source and limits: Equipping agents for the real world with Agent Skills, published October 16, 2025; checked October 7, 2026. This is vendor guidance, not a measurement of malicious-skill prevalence.
Commonset interpretation: Define what the review covers. An approval of the instruction text should not silently become an approval of a later-added script or changed dependency. Consider the whole package, the tools and data it can reach, and the environment in which it will run. Review is one control, not a guarantee of safe execution.
4. Native skill versioning exists; teams still need a version-selection policy
OpenAI's Skills API documents new-version uploads, a default-version pointer, a latest-version pointer, and explicit numeric version references. When a version is omitted, the default is used; latest tracks the newest upload.
Source and limits: OpenAI Skills API guide, Versioning and management, checked October 7, 2026. This describes that API surface, not every OpenAI product or every other provider.
Commonset interpretation: Use native controls where they meet the requirement. Decide which workflows may follow updates and which must stay on a reviewed revision. For a capability used through several platforms, also decide how each platform's revision maps to the organization's approved version. The useful question is whether the intended version reaches the intended users, not whether versioning exists at all.
Broader enterprise AI context
The findings below concern general AI or AI agents. They do not measure AI-skill inventories, version drift, approval coverage, or willingness to pay for governance software.
Gallup's indicator describes weighted research using its probability-based U.S. panel, covering employed adults. The linked publications do not state the May 2026 wave's sample size. McKinsey's 2026 online survey covered 1,719 respondents in 97 countries, fielded May 4–June 8, 2026, with results weighted by national contribution to global GDP. Both sources rely on respondent reports, rather than an independent audit of each organization's systems. See Gallup's survey methods and McKinsey's report methodology.
5. Occasional AI use and routine AI use are different measures
Gallup reports that, in May 2026, 52% of U.S. employees used AI at work at least a few times a year, 30% used it at least a few times a week, and 15% used it daily. These are overlapping frequency categories, not separate groups to add together.
Source: Organizational AI Adoption Jumps Six Points, July 20, 2026, read alongside the indicator's frequency definitions.
Commonset interpretation: Start a team inventory with recurring work, not a count of accounts or occasional experiments. Identify the tasks people repeat, the reusable instructions they rely on, and where those instructions are kept. Establish the workflow before deciding what governance it needs. These percentages cannot tell you how many employees use shared skills.
6. A communicated AI plan is not the same as an implemented control
Gallup's May 2026 indicator reports that 25% of U.S. employees said their organization had communicated a clear plan for integrating AI into existing practices.
Source and limits: Gallup Artificial Intelligence indicator, checked October 7, 2026. This measures employee perception of communication. It is not an audit showing that the remaining organizations lack policies, technical controls, or a strategy.
Commonset interpretation: Make the operational answer easy to find: where should an employee get an approved skill, who owns it, and what should they do when it fails or needs updating? A policy document and a discoverable, usable process are different deliverables. Do not subtract this figure from an adoption statistic and label the difference ungoverned AI.
7. Agent scaling differs by organizational size
McKinsey reports that 40% of respondents from organizations with at least US$1 billion in annual revenue said their organizations were scaling agents in one or more functions, compared with 22% from smaller organizations.
Source and limits: The state of AI in 2026: On the road to ROI, August 25, 2026. These are survey reports about agent scaling, not counts of deployed skills or a census of companies.
Commonset interpretation: Scope a governance discussion to an actual deployment. Ask which function is using agents, what those agents can do, which reusable instructions shape their work, and who can approve changes. Organization size alone is a poor substitute for finding a real owner, a recurring workflow, and a consequential problem.
8. Personal productivity and enterprise financial impact are separate outcomes
In the same McKinsey survey, 80% of respondents reported improved personal productivity from AI, while 37% reported a positive contribution to their organization's earnings before interest and taxes.
Source and limits: McKinsey's 2026 report. These are different self-reported outcomes. The figures do not establish that governance causes profitability, or that organizations without reported earnings impact receive no benefit.
Commonset interpretation: Evaluate reusable work at the workflow level. Choose an outcome such as correct completion, turnaround time, or rework, and examine what happens when another person uses the same capability. A registry entry, an approval, or a high usage count is not itself evidence of a business improvement.
Turning the evidence into a working practice
Choose one recurring workflow. Name the person responsible for the reusable capability, identify the source and exact revision, and define the conditions under which it should be used. Evaluate that revision on representative work before recommending it to other people.
Then test the handoff. Can a second person find the intended version and use it without relying on the author's memory? When a correction is approved, is there a defined way to update the relevant users or systems? When it fails, is there a record that connects the outcome to the version and operating environment?
These are proposed management practices, not a claim that the cited studies tested Commonset. Their value should be judged against the team's actual work and existing platform controls.
For implementation guidance, see Managing Claude Skills Across an Enterprise and Versioning and Comparing Changes. For the broader management model, see What Is an AI Skill Registry?.