We Analyzed 250 Enterprise AI Case Studies. Here’s Where AI Is Actually Paying Off
The Brief
Original research based on 250 enterprise AI cases, with case-level sources, evidence scoring, financial-ROI classification, failure analysis, validation checks, and a full research appendix. RESEARCH WINDOW January 1, 2024 – July 25, 2026 | 250 analytical cases | 209 organization keys | TheBriefScript Research Team Executive summary Enterprise AI is scaling faster than companies are […]
Why It Matters
The story matters because it changes how buyers, builders, or policymakers should read the Enterprise AI market.
Watch Next
Watch whether the signal becomes a budget, procurement, or platform decision in the next cycle.
Original research based on 250 enterprise AI cases, with case-level sources, evidence scoring, financial-ROI classification, failure analysis, validation checks, and a full research appendix.
| RESEARCH WINDOW January 1, 2024 – July 25, 2026 | 250 analytical cases | 209 organization keys | TheBriefScript Research Team |
Executive summary
Enterprise AI is scaling faster than companies are proving its financial return. Across 250 enterprise AI deployments, experiments, company disclosures, filings, government evaluations, and independent studies, 76% had reached scaled production and 90% published some quantified outcome. But only 29 cases, or 11.6%, explicitly disclosed money, ROI, payback, breakeven, or total cost of ownership in the primary published result.
Only 10 of 250 cases, or 4%, combined an observed operational result with explicit monetized or ROI evidence. That does not mean only 4% of enterprise AI creates value. It means the public evidence base is much stronger on time savings, throughput, quality, accuracy, and adoption than on realized financial return.
| CORE THESIS Enterprise AI is creating measurable operational value. The harder problem is proving where that value goes. |
Key numbers
| Metric | Result |
| Cases analyzed | 250 |
| Scaled production | 76.0% |
| Quantified outcomes | 90.0% |
| Strong + scaled + quantified | 47.6% |
| Explicit monetized / ROI evidence | 11.6% |
| Commercial proxy evidence | 8.0% |
| Any financial / commercial evidence | 19.6% |
| Observed + explicit monetized / ROI | 4.0% |
| Projected or estimated claims | 20.8% |
| Adoption-only cases | 7.6% |
| Failure or materially mixed cases | 6.0% |

Figure 1. Production scale versus public financial evidence in the 250-case dataset.
About this research
TheBriefScript reviewed 250 enterprise AI cases covering 209 organization keys between January 2024 and July 2026. Each case was scored against a 100-point evidence framework and classified by claim basis, financial evidence, production maturity, source group, use case, and industry. The full 250-case evidence index and validation notes are included in the appendices of this publication.
Research status: dataset frozen; duplicate audit, statistical analysis, organization sensitivity testing, and AI second-pass validation complete; independent human second review pending.
1. Methodology and scoring framework
We collected enterprise AI deployments published between January 1, 2024 and July 25, 2026. Eligible organizations generally had at least 500 employees, were major public or institutional entities, or operated deployments serving at least 1,000 users or customers.
The sample includes generative AI assistants, copilots, agents, coding tools, enterprise search, customer service, marketing, finance and risk, healthcare, manufacturing, supply chain, R&D, and other production AI workflows.
We excluded partnership announcements without implementation results, generic product launches, future plans without deployments, unnamed survey statistics, duplicate deployments, and cases without a traceable primary or credible source.
Evidence-quality scoring
| Criterion | Points | What it tests |
| Quantified result | 15 | Does the source publish a numerical outcome? |
| Baseline disclosure | 10 | Is there a usable comparator or pre-AI baseline? |
| Measurement period | 10 | Is the observation or evaluation window disclosed? |
| Deployment / sample size | 10 | Are users, tasks, records, calls, visits, documents, or similar denominators disclosed? |
| Methodology detail | 10 | Does the source explain how the result was calculated or evaluated? |
| Production maturity | 15 | Is the system in real production rather than only a demo or early pilot? |
| Business materiality | 10 | Is the outcome tied to a business, operational, clinical, risk, or financial metric? |
| Denominator / magnitude | 5 | Can the scale of the result be understood? |
| Adopter evidence | 10 | Is the adopter itself a direct evidence source or clearly identified? |
| Limitations disclosed | 5 | Does the source acknowledge caveats, uncertainty, mixed effects, or constraints? |
Evidence bands: Strong = 80–100; Moderate = 60–79; Limited = 40–59; Weak = below 40.
| IMPORTANT The evidence score measures transparency and evaluation quality. It is not an ROI score and does not independently verify that every published claim is true. |
Unit of analysis and duplicate control
The unit of analysis is an enterprise AI deployment, workflow, evaluation, or clearly bounded use case. A company could appear more than once only when the underlying workflow was materially different. For example, a bank’s fraud agent and its internal employee assistant can be retained as separate cases because they measure different tasks, populations, and business outcomes. This prevents the dataset from treating every announcement by a large organization as independent evidence.
Before the final freeze, 251 records were reviewed for duplication. One genuine duplicate was consolidated: two Klarna records described the same OpenAI-powered customer-service deployment. The retained case was expanded with later follow-up evidence, and the duplicate record was removed. The final analytical dataset contains 250 cases with no duplicate case IDs and no exact organization-platform-use-case duplicates.
We also tested whether organizations with multiple retained workflows could dominate the results. The 250 cases map to 209 organization keys. Twenty-two organizations contribute more than one case, the maximum for any single organization is seven, and the ten most represented organizations account for 15.6% of the dataset. The case-weighted average evidence score was 80.43 before adjudication, while the organization-balanced average was 80.20, a difference of only 0.23 points. Repeat organizations therefore do not materially drive the headline evidence score.
Source composition and sampling controls
The research was designed to avoid becoming a collection of vendor success stories. Vendor and partner libraries are still included because they often contain the clearest production-scale deployment details, but no single vendor ecosystem was allowed to exceed 10% of the final sample. Microsoft was the largest individual vendor ecosystem at 23 cases, or 9.2% of the dataset.
The final source mix contains 135 vendor or partner-library cases, 71 adopter-owned or financial disclosures, and 44 independent or verified research cases. Adopter-owned, regulatory, financial, and independent evidence together represent 46% of the final sample, above the study’s 38% minimum. Independent evidence was intentionally oversampled because formal experiments and mixed outcomes were important for testing whether production case-study claims survive more rigorous evaluation.
This balancing choice means the dataset should not be interpreted as a market-share sample of enterprise AI vendors. It is an evidence sample designed to compare how value is measured across production deployments, company disclosures, and formal evaluations.
How financial evidence was classified
Financial disclosure was coded separately from the 100-point evidence score. A case could score highly for transparency even when the reported result was projected rather than realized. Conversely, a deployment could be operating at large scale while publishing no financial result at all. Keeping these concepts separate prevents a polished case study from being mistaken for proven cash ROI.
For the publication analysis, we classified the primary published result into three financial-evidence levels. Monetized / ROI requires explicit currency, ROI, payback, breakeven, or total-cost-of-ownership language. Commercial proxy includes revenue, sales, pipeline, costs, losses, fraud exposure, or a similar business result without a direct monetized ROI statement. Non-financial includes time, throughput, quality, accuracy, adoption, satisfaction, and other operating outcomes without a financial or commercial result.
This is deliberately conservative. A statement such as ‘250,000 hours saved annually’ remains non-financial unless the source itself connects that capacity to money, payback, or another explicit commercial result. Likewise, $20 million of pipeline is classified as a commercial signal rather than realized revenue. The taxonomy describes what companies publicly disclose rather than converting every operational metric into a speculative dollar value.
Claim basis: observed, projected, controlled, adoption-only, and mixed
Each case was also classified by the basis of the business-value claim. Observed operational or financial cases report an outcome from a deployed workflow. Projected or estimated cases publish a modeled, annualized, expected, or forecast benefit. Controlled or independent cases use experiments, formal observational designs, blinded review, or similarly rigorous evaluation. Adoption-only cases show meaningful deployment or usage without a corresponding business-value result. Failure or mixed cases include reversals, statistically null results, material quality tradeoffs, legal exposure, safety concerns, or other evidence that complicates a simple success narrative.
This distinction is central to the analysis because a forecast can be transparent and well documented while still not being realized ROI. The evidence score asks how well the claim is disclosed. The claim-basis taxonomy asks what kind of claim it is. The financial taxonomy asks whether the public result connects the claim to money or a commercial outcome.
Validation, adjudication, and sensitivity checks
After the 250-case dataset was frozen, a stratified 25-case sample was recoded in a separate AI second-pass audit using the same rubric. This was a validation consistency exercise, not an independent human review. The second pass produced 72% evidence-band agreement, a Pearson score correlation of 0.81, and quadratic weighted band kappa of 0.67. The mean absolute score difference was 6.36 points, with 80% of cases within plus or minus 10 points.
The validation audit was most sensitive on limitations disclosure, measurement period, baseline detail, and denominator disclosure. These are exactly the criteria where case studies often imply context without explicitly stating it. The primary coding sometimes gave partial credit where the second pass required more explicit source language. Adopter evidence moved in the opposite direction in several cases because the second pass gave more credit when named organizations clearly supplied internal operational data.
Two discrepancies were factual extraction issues rather than rubric judgment. London Stock Exchange Group had explicit OpenAI-reported delivery outcomes that were missed in the original extraction, including product release cycles moving from three to six months to roughly two weeks. Armanino’s source also contained more explicit workflow baselines, deployment scale, and estimated benefits than the primary row captured. Those two cases were corrected in the adjudicated dataset. Rubric-sensitive disagreements were left unchanged pending an independent reviewer.
The corrections barely moved the headline dataset. The average evidence score increased by 0.20 points, the median remained 83, the Strong band gained one case, and the quantified-outcome rate moved from 89.6% to 90.0%. This sensitivity check shows that the central findings are not dependent on one or two disputed records.
What this research can and cannot establish
The study can show what kinds of enterprise AI value are being publicly documented, where the evidence is strongest, which workflows are easiest to measure, how often financial evidence is disclosed, and how production case studies differ from independent evaluations. It cannot estimate the true success rate of all enterprise AI projects because the underlying evidence is not a random sample and public success stories are more visible than quiet failures.
For the same reason, the 90% quantified-outcome rate is not a claim that 90% of enterprise AI projects succeed. It means 225 of the 250 selected and source-verified cases contained at least one numerical outcome after adjudication. Those numbers range from operational improvements and quality metrics to adoption measures, projections, controlled experimental effects, and mixed or negative findings.
2. Where enterprise AI is actually paying off
The strongest pattern was not tied to one model or vendor. It was tied to the structure of the work. AI performed best when the workflow had a clear unit of work, a measurable baseline, enough repetition to generate telemetry, and an output that could be checked against an objective standard.
| Use-case family | Cases | Strong + scaled + quantified | Explicit monetized / ROI |
| Knowledge work & analytics | 62 | 43.5% | 6.5% |
| Software, IT & security | 54 | 57.4% | 13.0% |
| Marketing, sales & content | 33 | 57.6% | 24.2% |
| Customer service & support | 29 | 55.2% | 20.7% |
| Healthcare & clinical | 15 | 26.7% | 6.7% |
| Finance, risk & compliance | 18 | 44.4% | 11.1% |
| Industrial, supply chain & operations | 15 | 46.7% | 6.7% |
| Research & R&D | 12 | 16.7% | 0.0% |

Figure 2. Share of cases with Strong evidence, scaled production, and a quantified outcome.
3. Software, IT, and security have the strongest evidence base
Software, IT, and security accounted for 54 cases and produced one of the strongest combinations of evidence quality and production maturity. About 57.4% of cases met the stricter threshold of Strong evidence, scaled production, and a quantified outcome.
The advantage is measurement. Engineering teams can track pull requests, acceptance rates, defects, cycle times, lines modernized, incidents closed, alerts reviewed, or support tickets resolved.
Booking.com reported 97% correct ticket assignment, a 33% reduction in resolution time, quality-review coverage rising from 7% to 100%, and a 30% increase in quality scores over six months.
IBM reported that an internal security deployment analyzed 874 applications in 24 hours, identified 32% more vulnerabilities, and reduced low-priority alerts by 67%.
Thomson Reuters reported modernizing 1.5 million lines of code per month, increasing modernization velocity fourfold and cutting costs by 30%. The financial attribution is not pure AI because the migration itself also changed infrastructure costs, but the workflow is unusually measurable.
Related TBS analysis: Why 62% pilot but only 23% deploy
4. Marketing and sales publish the most commercial evidence
Marketing, sales, and content produced a similar strong-at-scale rate of 57.6%, but stood out for another reason: 39.4% of cases contained either explicit monetized evidence or a commercial proxy such as sales, pipeline, conversion, or cost. That is the highest concentration among the major use-case families in the study.
Salesforce reported that its internal Agentforce deployment generated more than 30,000 leads in 2025, increased lead conversion 1.8 times, and was associated with a $20 million pipeline lift.
Mattel reported an estimated $1 million in cost savings and a dramatic increase in customer-feedback processing capacity. Emccamp reported cutting one landing-page production cost from R$8,000 to R$50 while reducing delivery time from three weeks to four hours.
These cases are commercially meaningful, but they also show why financial claims need careful classification. Pipeline is not revenue. Estimated savings are not audited cash savings. Faster content creation does not automatically mean incremental profit.
5. Customer service is both a value cluster and a failure cluster
Customer support is one of the most mature areas of enterprise AI. More than half of the cases in this category met the Strong, scaled, and quantified threshold.
Chime reported an 18-second reduction in average call handling time, more than 250,000 hours saved annually, an estimated $700,000 in annual efficiency gains, and a five-point increase in service NPS.
Nubank reported an AI assistant handling more than two million chats per month, resolving roughly 55% of Tier 1 inquiries and reducing response time by 70%.
But customer service also contains some of the clearest evidence that autonomy has limits. McDonald’s ended a drive-through AI ordering test spanning more than 100 restaurants. DPD disabled an AI chatbot component after a highly visible customer interaction. Air Canada was held responsible for inaccurate information supplied by its chatbot.
The conclusion is not that customer-service AI does not work. It is that routine resolution and information retrieval are much easier to automate than accountability, empathy, policy interpretation, and edge cases.
6. Industrial AI performs well when the unit of work is physical and measurable
Industrial and manufacturing cases had the highest industry-level share of Strong, scaled, and quantified evidence in the study at 72.2%. These deployments often operate against hard metrics: inspection time, energy consumption, engineering cycle time, production planning, defect detection, procurement, inventory, or plant penalties.
ABB reported an industrial energy deployment that reduced energy costs by 1.5%, day-ahead penalties by 60%, and intraday penalties by 80%, with ROI achieved within one year.
Mercedes-Benz reported a 20% reduction in energy consumption from AI-controlled paint-shop process engineering. BMW reported digital-factory infrastructure spanning more than 30 production sites and projected up to a 30% reduction in production-planning costs.
7. The knowledge-work ROI gap
Knowledge work and analytics was the largest use-case family in the dataset with 62 cases. It included Microsoft 365 Copilot deployments, ChatGPT Enterprise rollouts, internal search assistants, meeting tools, document drafting systems, research assistants, and general-purpose productivity copilots.
Yet only 6.5% of knowledge-work cases disclosed explicit monetized or ROI evidence. The dominant evidence was instead expressed in hours saved.
MIXI estimated 17,600 work hours saved monthly. Pennsylvania’s government pilot reported average self-assessed savings of 95 minutes on days employees used ChatGPT. Lloyds reported average Copilot savings of 46 minutes daily in an internal survey. Kantar reported average savings of 1.8 hours per week among enabled users.
An employee saving 45 minutes does not mean payroll expense falls by 45 minutes. The time may become additional capacity, better work, shorter queues, more customer conversations, or simply unused capacity. This is why “hours saved × salary” should be treated as a capacity estimate unless the company demonstrates what happened to the released time.
Related TBS analysis: The AI Value Crisis: Why Enterprises Are Struggling to Monetize AI in 2026
8. Why healthcare looks different
Healthcare produced some of the strongest evidence quality in the dataset, but not because it had the most financial disclosures. It had better study designs.
Randomized trials, stepped-wedge trials, interrupted time-series analysis, difference-in-differences designs, blinded clinical review, and large encounter datasets appeared far more often than in ordinary enterprise case studies.
At Penda Health, a real-world deployment study covering 39,849 visits found 16% fewer diagnostic errors and 13% fewer treatment errors among clinicians with AI support. Independent physicians reviewed 5,666 randomly selected visits.
A UCSF study covering more than 1.2 million encounters found modest productivity gains associated with ambient AI adoption. A UCLA randomized evaluation comparing two ambient AI tools found that one significantly reduced time in notes while another did not produce a statistically significant time reduction.
Healthcare therefore offers some of the study’s strongest causal evidence while simultaneously demonstrating why simple “AI saves X%” headlines can be misleading.
9. How much financial ROI is actually public?
We classified financial evidence conservatively using the primary published result. A case was classified as Monetized / ROI only when the public result explicitly included currency, ROI, payback, breakeven, or total cost of ownership language. A Commercial proxy included revenue, sales, pipeline, costs, fraud losses, or a similar commercial result without a direct monetized ROI disclosure.
| Financial evidence level | Cases | Share of sample | Average evidence score |
| Monetized / ROI | 29 | 11.6% | 88.0 |
| Commercial proxy | 20 | 8.0% | 83.7 |
| Non-financial | 201 | 80.4% | 79.3 |
The 29 monetized cases should not all be interpreted as realized cash ROI. Fourteen sit inside the dataset’s projected or estimated category. Others report annualized capacity values, potential savings, identified financial risk, legal liability, or payback calculations that use assumptions not fully disclosed publicly.
Only 10 cases – 4% of the total dataset – combined the observed-value category with explicit monetized or ROI disclosure. That is the closest public proxy in this dataset to “enterprise AI has produced a financial result that the company is willing to quantify.” It is still not equivalent to audited causal ROI.
| THE ROI GAP Large deployment numbers demonstrate organizational commitment. They do not, by themselves, demonstrate productivity, cost savings, revenue growth, or return on investment. |
Why 90% quantified does not mean 90% successful
The most easily misunderstood number in the study is the 90% quantified-outcome rate. On its own, it sounds like a success statistic. It is not. A quantified outcome simply means the source published a number tied to the deployment or evaluation. That number could describe faster work, higher quality, more users, an expected saving, a controlled-study effect, a legal loss, or a failed experiment.
After adjudication, 135 cases, or 54.0%, fall into observed operational or financial value. Another 52 cases, or 20.8%, are projected or estimated. Twenty-nine cases, 11.6%, come from controlled or independent evaluations. Nineteen cases, 7.6%, are adoption-only, and 15 cases, 6.0%, are failure or materially mixed.
The resulting picture is more useful than a binary success rate. More than half of the sample contains observed operating outcomes. Roughly one-fifth contains projections that may or may not be realized. A meaningful research subset shows controlled effects, often at limited scale. A smaller group shows adoption without a business result, and another group documents boundaries, reversals, or mixed performance.
| Claim basis | Cases | Share | Average evidence score | Interpretation |
| Observed operational / financial | 135 | 54.0% | 80.95 | Measured outcome from a deployed workflow |
| Projected / estimated | 52 | 20.8% | 83.83 | Expected, modeled, annualized, or forecast value |
| Controlled / independent | 29 | 11.6% | 88.07 | Formal experiment or independent evaluation |
| Adoption only | 19 | 7.6% | 60.47 | Meaningful deployment without business-value evidence |
| Failure / mixed | 15 | 6.0% | 77.80 | Null, negative, reversed, risky, or materially mixed outcome |
Controlled evidence is rigorous, but usually not scaled
Controlled and independent cases have the highest average evidence score in the study, 88.07, but only 17.2% are in scaled production. That is not a contradiction. Rigorous measurement is often easiest in a bounded environment where researchers can define the sample, control the intervention, and observe the same task across treated and comparison groups.
Production deployments face the opposite tradeoff. Companies can show that thousands of employees use a system, millions of calls are processed, or a workflow operates across dozens of sites, but they often cannot isolate what would have happened without the AI. Adoption changes over time, workers learn, processes are redesigned, demand shifts, and other software changes at the same time. The result is strong evidence of scale with weaker evidence of causality.
This is why the research should not force controlled studies and production case studies into a single hierarchy. A randomized trial can tell us whether a tool caused an effect under defined conditions. A large enterprise deployment can tell us whether the workflow survives real data, security requirements, integration complexity, user behavior, and operational load. The most persuasive future evidence will combine both.
The financial evidence ladder
Financial evidence becomes much thinner as the standard gets stricter. Seventy-six percent of cases reached scaled production. Forty-seven point six percent combined Strong evidence, scaled production, and a quantified result. Nineteen point six percent published either financial evidence or a commercial proxy. Eleven point six percent explicitly disclosed money, ROI, payback, breakeven, or TCO. Only 4.0% combined explicit monetized evidence with the observed-value class.
This is the central enterprise AI ROI gap. Companies are increasingly able to demonstrate that AI is deployed, used, and changing an operating metric. Far fewer publicly show that the change flowed through to realized savings, incremental revenue, avoided losses, or a measured return on investment.
The 29 monetized or ROI cases also need interpretation. Fourteen belong to the projected or estimated class. Some use annualized capacity, expected future savings, or payback calculations. Three are failure or mixed cases where money appears because the source quantifies damages, exposure, or another adverse financial consequence. A monetary number therefore makes a case more commercially legible, but it does not automatically make the case a success.
Projected ROI is more common than observed monetized ROI
The difference is clearest when monetized disclosure is cross-tabulated with claim basis. Fourteen of 52 projected or estimated cases, 26.9%, disclose explicit monetized or ROI evidence. Only 10 of 135 observed cases, 7.4%, do the same. Controlled studies are even less likely to publish money: two of 29, or 6.9%. None of the 19 adoption-only cases publishes explicit monetized ROI.
This does not mean projections are useless. Forecasts can be necessary for capital allocation, budgeting, and deciding whether to move from pilot to production. But they answer a different question from realized operating evidence. A projected $5 million benefit is a business case. An observed reduction in processing time is an operating result. A verified reduction in total cost after implementation is closer to realized ROI. The study keeps those layers separate rather than collapsing them into one number.
| Value status | Cases | Monetized / ROI cases | Rate |
| Observed | 135 | 10 | 7.4% |
| Projected / estimated | 52 | 14 | 26.9% |
| Controlled / independent | 29 | 2 | 6.9% |
| Adoption only | 19 | 0 | 0.0% |
| Failure / mixed | 15 | 3 | 20.0% |
Adoption without ROI is a real research finding
Nineteen cases in the final dataset show meaningful adoption without a corresponding business-value result. These are not mistakes in the sample. They are intentionally retained because removing them would create measurement cherry-picking. Enterprise AI is often announced through user counts, licenses, prompts, agents built, or access expanded long before the adopter publishes what changed in the business.
The adoption-only group has an average evidence score of about 60.5 and contains no Strong cases. That gap is instructive. A company can disclose tens of thousands of enabled employees and still leave basic questions unanswered: Which workflows changed? How much time was released? Did quality improve? Did cost fall? Did customers notice? Did revenue or risk move?
The practical implication is that reach should be reported as reach. A 35,000-seat Copilot deployment can be strategically important, but the seat count should not be presented as productivity evidence. Likewise, high weekly active usage can demonstrate that employees found reasons to use the system, but it does not show whether those uses produced incremental enterprise value.
10. What the strongest enterprise AI deployments have in common
1. They define the unit of work
Strong cases measure claims, calls, reports, pull requests, queries, documents, transactions, inspections, meetings, tickets, or another concrete operational unit. “Employees are more productive” is weak evidence. A before-and-after unit with a denominator is much stronger.
2. They provide a baseline
Nearly all Strong evidence cases disclose some form of baseline. A percentage without a starting point is difficult to interpret. “50% faster” means much more when the source explains whether the task previously took four minutes, four hours, or four weeks.
3. They measure long enough to distinguish a demo from a system
Measurement-period disclosure appears far more consistently in Strong cases. Results are more useful when a source states whether the effect came from a one-hour task test, a six-week pilot, six months of production telemetry, or a full financial year.
4. They explain how the number was produced
Methodology is one of the largest disclosure gaps between Strong and non-Strong evidence. The best examples explain the calculation, comparison group, sample, instrumentation, or evaluation design rather than publishing a headline percentage in isolation.
5. They keep humans where the task demands judgment
The strongest systems often automate part of a workflow rather than the entire job. Fraud agents require approvals, clinical copilots assist clinicians, and customer-service systems route difficult cases to people.
11. Where enterprise AI breaks down
Fifteen cases in the final dataset were classified as failure or materially mixed outcomes. That number should not be interpreted as a 6% enterprise AI failure rate. Failed projects are much less likely to be published than successful case studies, and the sample was not designed to estimate the population failure rate.
1. Open-ended judgment is harder than narrow task execution
A controlled UK AI Security Institute study found overall improvements in work quality and speed, but an open-ended planning task produced no statistically significant improvement. An Irish public-sector experiment found that AI improved document-task quality and speed while reducing quality on a data-analysis task.
2. High-stakes interpretation creates asymmetric risk
Cando Rail paused an internal safety chatbot after it inconsistently interpreted operating rules. Air Canada faced legal liability after its chatbot supplied incorrect fare-policy information. New York City’s public business chatbot was found giving inaccurate or potentially unlawful regulatory guidance.
3. Autonomous customer interaction exposes failure publicly

Figure 3. Enterprise AI performs best in structured, lower-stakes work. Human review becomes more important as tasks become more open-ended and higher stakes.
The matrix is a synthesis of the failure and mixed-outcome cases in the dataset. The most persistent risk patterns appear in policy interpretation, high-stakes guidance, safety-sensitive decisions, and autonomous customer-facing edge cases.
Internal AI errors can be corrected quietly. Customer-facing errors can become reputational, legal, or regulatory events. The transition from agent-assist to autonomous agent should therefore be treated as a governance threshold, not merely a product feature.
Related TBS analysis: Enterprise AI ROI 2026: Why 80% of Projects Fail
12. Vendor stories show scale. Independent studies show how the measurement works.
| Evidence source | Cases | Average evidence score | Scaled production |
| Vendor/partner library | 135 | 80.3 | 89.6% |
| Adopter-owned/financial | 71 | 78.9 | 85.9% |
| Independent research | 44 | 84.6 | 18.2% |
Independent research had the highest average evidence score because it more consistently disclosed samples, baselines, methods, observation periods, and limitations. But independent studies were much less likely to represent full production scale.
Vendor stories showed the opposite pattern: dramatically more production evidence, but less consistent methodological disclosure. Neither source type is sufficient on its own. Independent studies answer, “Does this effect survive a better evaluation?” Production case studies answer, “Can this system actually operate inside a large enterprise?” The strongest enterprise AI evidence eventually needs both.

Figure 4. Independent studies explain the effect more clearly, while vendor and adopter cases provide much more evidence of scaled production.
This source split is one of the most important findings in the study. Independent evidence explains how the effect was measured. Enterprise case studies show whether the workflow can actually run at scale.
13. What this means for enterprise AI buyers
The most useful question for an AI buyer is no longer, “Does generative AI increase productivity?” The evidence is sufficient to say that it can. The better questions are:
- For which workflow?
- Against what baseline?
- For which workers?
- Over what period?
- At what quality level?
- At what implementation cost?
- What happened to the time or capacity released?
- Did the gain persist after the novelty period?
- Was revenue incremental or merely associated?
- Would the result survive a control group?
These questions turn an AI business case from a model demo into an operating decision.
Industry patterns
| Industry cluster | Cases | Strong + scaled + quantified | Measured non-projected | Monetized / ROI |
| Technology & software | 68 | 58.8% | 82.4% | 8.8% |
| Financial services | 62 | 50.0% | 62.9% | 12.9% |
| Consumer, retail & ecommerce | 22 | 22.7% | 54.5% | 9.1% |
| Telecom, logistics & travel | 20 | 70.0% | 70.0% | 20.0% |
| Industrial & manufacturing | 18 | 72.2% | 72.2% | 16.7% |
| Healthcare & life sciences | 17 | 29.4% | 88.2% | 5.9% |
| Professional services | 16 | 25.0% | 56.2% | 0.0% |
| Public sector & education | 12 | 8.3% | 91.7% | 16.7% |
| Other | 15 | 40.0% | 66.7% | 20.0% |
What the industry data says
Industry-level results reinforce the same measurement pattern. Industrial and manufacturing organizations have the highest share of Strong, scaled, and quantified cases at 72.2%. Telecom, logistics, and travel follow at 70.0%. Technology and software are lower at 58.8% but contribute the largest absolute number of high-quality scaled cases because the category itself is much larger.
These sectors share something important: many workflows have visible units. Industrial systems can measure energy, defects, planning cycles, maintenance events, penalties, or production throughput. Telecom and logistics can measure resolution, routing, network work, handling time, or operational volume. Software companies can measure pull requests, defects, code accepted, incidents closed, and release velocity.
Healthcare and the public sector look different. Healthcare and life sciences show an 88.2% measured non-projected rate, while public sector and education reach 91.7%. Yet their Strong + scaled + quantified shares are only 29.4% and 8.3%, respectively. Those categories contain more formal evaluations and trials, but less full-scale production evidence. The data again shows that methodological rigor and production maturity are different dimensions.
Financial services are highly scaled at 90.3% of cases, but only 12.9% publish explicit monetized or ROI evidence. This is notable because banks and insurers operate close to measurable financial outcomes. Even there, the public evidence more often describes time saved, service quality, fraud detection, throughput, or employee adoption than a clean causal ROI calculation.
What separates Strong evidence from the rest
Because the evidence score is constructed from disclosure criteria, comparisons between Strong and non-Strong cases are descriptive rather than causal. Still, the gaps show exactly what information is missing from weaker enterprise AI claims.
All Strong cases contain a quantified outcome, compared with 74.5% of non-Strong cases. Baseline disclosure appears in 98.7% of Strong cases versus 75.5% outside the Strong band. Measurement-period disclosure is 93.4% versus 64.3%. Methodology detail shows the largest gap: 97.4% of Strong cases publish enough information to understand how the result was produced, compared with 61.2% of non-Strong cases.
Production scale, by contrast, barely separates the groups. Seventy-eight point three percent of Strong cases are scaled compared with 72.4% of non-Strong cases, a gap of only 5.8 percentage points. That is a crucial finding for buyers. Scale is relatively easy to demonstrate. Transparent measurement is much harder.
| Disclosure criterion | Strong cases | Non-Strong cases | Gap |
| Quantified outcome | 100.0% | 74.5% | 25.5 pp |
| Baseline | 98.7% | 75.5% | 23.2 pp |
| Measurement period | 93.4% | 64.3% | 29.1 pp |
| Deployment / sample size | 100.0% | 82.7% | 17.3 pp |
| Methodology detail | 97.4% | 61.2% | 36.1 pp |
| Scaled production | 78.3% | 72.4% | 5.8 pp |
The lesson is not that every enterprise case study needs to become an academic paper. It is that basic disclosure dramatically improves interpretability. State the baseline. State the period. State the denominator. Explain whether the number comes from telemetry, a survey, an experiment, an estimate, or an executive calculation. Say whether the system is in a pilot, limited deployment, or scaled production. These details turn an impressive percentage into evidence that another buyer can evaluate.
The dataset is not being driven by a handful of repeat companies
Large technology vendors and banks appear repeatedly in enterprise AI research because they operate many distinct workflows. That creates a legitimate concentration concern: if one organization produces several high-scoring cases, could it artificially lift the dataset? The organization-balanced sensitivity test suggests the answer is no.
IBM is the most represented organization with seven cases, followed by Salesforce with five. Vodafone, Lloyds Banking Group, DBS, and JPMorganChase each contribute four. Yet the case-weighted and organization-balanced averages differ by only 0.23 points in the frozen dataset. The headline evidence score therefore remains essentially unchanged when each organization is given equal aggregate weight regardless of how many workflows it contributes.
The strongest enterprise AI evidence is narrow before it is broad
Across the dataset, the clearest results come from bounded problems: summarize this call, classify this ticket, modernize this code, retrieve this policy, draft this rejection note, inspect this record, route this request, identify this vulnerability, or optimize this energy decision. The task has a start, an end, and an observable output.
General-purpose knowledge work is harder because the unit of value is less stable. An employee can use an assistant for writing, searching, brainstorming, analysis, meetings, email, and planning in the same hour. A survey can estimate time saved, but converting that time into enterprise value requires another layer of evidence about capacity, quality, output, or financial impact.
That helps explain why the enterprise AI story is becoming less about model capability and more about workflow design. The model may be general, but the strongest business case is usually specific. Enterprises appear to create the clearest value when they decide exactly what unit of work AI is supposed to change, instrument that unit, and keep the human decision boundary visible.
A practical evidence standard for AI buyers
The dataset suggests a simple test for evaluating a vendor case study or an internal pilot. First, ask what changed. Second, ask what the baseline was. Third, ask how long the result was observed. Fourth, ask how many users, tasks, transactions, or records were included. Fifth, ask whether the number was measured, surveyed, annualized, projected, or experimentally estimated. Sixth, ask what happened after the system reached production.
For financial claims, add three more questions. Is the number realized or forecast? Is it cost removed, capacity created, revenue associated, pipeline generated, loss avoided, or risk identified? And can the organization explain the attribution path from the AI intervention to the financial outcome? These distinctions are essential because enterprise AI can create genuine operating value without immediately producing a clean accounting line.
This standard also protects against the opposite mistake: dismissing useful AI simply because a company cannot yet prove audited ROI. For some workflows, a quality improvement, risk reduction, shorter queue, faster investigation, or higher throughput may be strategically valuable before finance can isolate the cash effect. The right standard is not ‘show dollars or the project failed.’ It is ‘label the evidence honestly and measure the next layer of value.’
14. Study limitations
This study measures the quality of publicly available enterprise AI evidence. It does not measure the true success rate of every enterprise AI deployment.
Successful deployments are more likely to become public case studies than unsuccessful ones. Vendor libraries are intentionally represented because they contain valuable production-scale evidence, but the publication incentives of vendors and adopters create unavoidable selection bias.
Financial results are not standardized across currencies, accounting definitions, time periods, attribution methods, or definitions of value. “$1 million saved,” “$1 million of capacity created,” and “$1 million of pipeline generated” are not economically equivalent.
Some organizations appear more than once where they operate genuinely distinct deployments. The organization-concentration sensitivity analysis found that the case-weighted and organization-balanced evidence-score averages remained nearly identical.
A 25-case AI validation recode found 72% evidence-band agreement and identified two factual extraction issues that were corrected in an adjudicated sensitivity dataset. Other scoring disagreements remain pending independent review.
For this reason, the article should describe the dataset as a structured review of published evidence, not a random sample of all enterprise AI projects.
The study is also constrained by what organizations choose to publish. Companies rarely disclose failed internal pilots, security incidents, abandoned experiments, or negative ROI calculations with the same detail used in success stories. The failure and mixed-outcome cases are therefore analytically useful but should not be used to estimate an overall enterprise AI failure rate.
Vendor case studies can contain accurate operational data while still reflecting selection incentives. Independent studies can provide stronger causal designs while testing narrower tasks, shorter periods, or less mature systems. Neither source type should be treated as a universal gold standard.
Finally, the independent human second review of the 25-case validation sample remains pending. The AI second-pass audit and two-case adjudication demonstrate that the central statistics are stable, but they do not replace an independent reviewer. The publication therefore reports the validation status explicitly rather than presenting the coding process as fully independently replicated.
15. FAQs
1. Does enterprise AI actually deliver ROI?
Yes, some deployments publish clear cost savings, ROI, payback, revenue, or commercial outcomes. But explicit monetized or ROI evidence appeared in only 11.6% of the 250 cases reviewed, despite 76% being in scaled production.
2. Which enterprise AI use cases have the strongest evidence?
Software development, IT, cybersecurity, marketing, sales, customer support, and narrowly defined industrial workflows have the strongest combination of production scale, quantification, and evidence quality in this dataset.
3. Why is knowledge-worker AI harder to measure?
Most general productivity tools create time or capacity rather than directly removing cost. Translating 45 minutes saved into financial value requires showing what happened to that capacity afterward.
4. Are vendor AI case studies reliable?
They are valuable evidence of production scale and operational implementation, but their public methodology is often incomplete. We treat them as evidence, not independent verification.
5. Why do independent studies score higher?
Independent evaluations more often disclose samples, controls, observation periods, methods, statistical uncertainty, and limitations. Their weakness is that they are less likely to represent companywide production systems.
6. What is the biggest enterprise AI measurement mistake?
Treating adoption as ROI. Users, licenses, prompts, agents, or models deployed show reach. They do not prove that cost fell, revenue increased, quality improved, or risk declined.
7. What should companies measure before deploying AI?
Define one operating outcome before rollout: handling time, cycle time, errors, cost per transaction, conversion, claims processed, engineering throughput, revenue, risk, or another measurable business metric. Then document the baseline and measure the same metric after deployment.
Conclusion
Enterprise AI is past the stage where the only evidence is demos and experiments. Across 250 cases, we found widespread production deployment and substantial evidence of faster work, greater throughput, better information access, improved quality, and automated routine tasks.
But the financial evidence has not caught up with the deployment curve. Only 11.6% of cases explicitly disclosed money, ROI, payback, breakeven, or TCO. Only 4% combined that kind of disclosure with our observed-value classification.
The most convincing enterprise AI deployments share a common structure: a narrow workflow, a measurable unit of work, a clear baseline, production telemetry, and a result tied to something the business already cares about.
| THEBRIEFSCRIPT VIEW Enterprise AI is creating value. The harder problem is proving where that value goes. |
Research appendix
The following appendices are part of the research publication rather than a separate supplement. They expose the case-level evidence, scoring logic, sampling controls, and validation process behind the headline findings so readers can inspect how the analysis was built.
The public case index is intentionally compact. It shows the organization, industry, AI workflow, primary published result, evidence score and band, and a direct source link. The underlying research workbook retains the full row-level scoring components, caveats, classifications, duplicate audit, and validation notes.
Appendix A. 250-case evidence index
This index includes every analytical case in the frozen dataset. It is intentionally compact: the full row-level scoring components, caveats, and research notes remain preserved in the statistical workbook used to build this report.
| ID | Company | Industry | AI / workflow | Primary published result | Score / band | Source |
| TBS-P001 | CDW | Technology services / retail | Microsoft 365 Copilot – Employee productivity, meetings, email and content creation | 10,000-user rollout; 85% reported higher productivity, 77% faster task completion and 88% better work quality. | 63 / Moderate | Source |
| TBS-P002 | British Columbia Investment Management Corporation (BCI) | Asset management | Microsoft 365 Copilot; Azure OpenAI Service – Employee productivity, audit reporting and survey analysis | 84% of initial users reported 10–20% productivity gains; 2,300+ person-hours saved; audit-report effort fell 30%. | 90 / Strong | Source |
| TBS-P003 | Hitachi | Technology and industrials | GitHub Copilot – Software development and code generation | 83% completed tasks faster; average coding gains of 10–20%, up to 30%; code-generation rate rose from 78% to 99% in one validation application. | 100 / Strong | Source |
| TBS-P004 | JCB | Payments | Microsoft 365 Copilot – Meetings, internal search, email, translation and coding | 83% average monthly usage; roughly six hours saved per employee per month across the top five use cases. | 78 / Moderate | Source |
| TBS-P005 | Insight Enterprises | Technology services | Microsoft 365 Copilot – Employee productivity, RFP preparation and business documents | Usage among Canadian users rose from 43% to 93% in three months; 4,600 staff had access globally; RFP response time improved by about 50%. | 92 / Strong | Source |
| TBS-P006 | Personio | HR software | Amazon Bedrock – Customer-facing HR assistant and workflow automation | 70% of routine HR requests automated; reporting time fell 21%; customer teams reported an 18% productivity improvement. | 58 / Limited | Source |
| TBS-P007 | Boomi | Enterprise software | Amazon Q Developer – Code generation, documentation, testing and security scans | 20% engineering productivity increase; 20% of deployed code generated by the tool; rollout expanded from 20 pilot developers to 445 developers. | 90 / Strong | Source |
| TBS-P008 | Genpact | Professional services | Amazon Bedrock; Amazon Titan – Background-check validation and employee AI playground | 75% productivity increase, up to 90% reduction in manual verification and 80% validation accuracy; 2 million interactions within nine months. | 70 / Moderate | Source |
| TBS-P009 | The Depository Trust & Clearing Corporation (DTCC) | Financial-market infrastructure | Amazon Q Developer – Software development, testing and debugging | 40% developer-throughput increase; a ten-hour task fell to six hours; code defects declined 30% during the proof of concept. | 90 / Strong | Source |
| TBS-P010 | PivotRoots | Digital marketing and analytics services | Gemini; BigQuery – Conversational client data analysis | Client query turnaround fell from 24–48 hours to under two minutes; project analysis costs declined 15–20%, and up to 80% for some larger clients. | 65 / Moderate | Source |
| TBS-P011 | Starling Bank | Banking | Gemini; Google Cloud AI – Customer help centre, scam intelligence and banking assistance | 50% reduction in customer-service referrals and 8,000 hours saved per month. | 60 / Moderate | Source |
| TBS-P012 | Toolstation | Retail | Gemini Enterprise for Customer Experience / AI Commerce Search – Ecommerce product search | No-result searches fell from 2% to 0.1%; click-through rate rose 10%; search-based revenue increased 5.5%. | 60 / Moderate | Source |
| TBS-P013 | Mantel | Technology consulting | Gemini Enterprise; NotebookLM – Statements of work, project-history summaries and research production | Manual work fell 20–75% in a statement-of-work trial; a 30-page research paper fell from five months to two months. | 70 / Moderate | Source |
| TBS-P014 | GitLab | Enterprise software | Claude for Work / Claude Enterprise – Content, data analysis, RFP responses and sales workflows | Pilot participants reported 25–50% productivity gains and 98% satisfaction. | 53 / Limited | Source |
| TBS-P015 | TRY | Advertising and communications | Claude Enterprise – Creative, strategy and proposal workflows | 30% reduction in routine-task time, 40% faster proposal development and 50+ implemented use cases across 400+ professionals. | 65 / Moderate | Source |
| TBS-P016 | IG Group | Financial services | Claude for Work – Analytics, query handling, content and operational workflows | Analysts saved 70 hours weekly; payback was reported in under three months; some workflows saw 100% productivity improvements. | 70 / Moderate | Source |
| TBS-P017 | Block | Financial technology | Claude in Databricks; codename goose – Internal agent for data, SQL and workflow automation | 75% of engineers reportedly saved 8–10+ hours each week; engagement grew 40–50% weekly and adoption doubled in one month. | 72 / Moderate | Source |
| TBS-P018 | BBVA | Banking | ChatGPT Enterprise – Enterprise knowledge work and internal assistants | Around 100,000 employees deployed; roughly three hours saved per employee weekly; one assistant cut query handling from 7.5 minutes to about one minute. | 87 / Strong | Source |
| TBS-P019 | BNY | Financial services | OpenAI models; internal Eliza platform – Enterprise agents, legal review and knowledge workflows | Legal review time fell 75%, from four hours to one, across 3,000+ annual vendor agreements; 125+ AI tools were in production. | 82 / Strong | Source |
| TBS-P020 | Cisco | Technology | Codex – Build optimization, defect remediation and software migrations | Build times fell about 20%, saving 1,500+ engineering hours monthly; defect-resolution throughput improved 10–15×. | 82 / Strong | Source |
| TBS-P021 | London Stock Exchange Group (LSEG) | Financial data and markets | OpenAI models; ChatGPT Enterprise – Research, product development and trusted data workflows | LSEG reported reducing product-release cycles from 3–6 months to about 2 weeks and customer delivery timelines to about 4 weeks from request to production; thousands of employees were enabled within weeks. | 73 / Moderate | Source |
| TBS-P022 | IBM CIO Organization | Enterprise IT | watsonx Assistant for Z – Mainframe onboarding, incident resolution and patching | 300+ users; learning time fell 8%, incident resolution improved 10% and Db2 patching time fell 50%. | 90 / Strong | Source |
| TBS-P023 | IBM Software Team | Software development | watsonx Code Assistant – Code explanation, documentation, generation and testing | 107 teams reported 56% average time savings on code explanation; 153 teams reported 59% on documentation; selected tasks showed 90%+ savings. | 85 / Strong | Source |
| TBS-P024 | IBM People Analytics | Human resources analytics | watsonx BI – Natural-language workforce analytics | The six-month pilot reportedly changed report delivery from days to rapid answers, but no exact time, cost or productivity figure is published. | 45 / Limited | Source |
| TBS-P025 | IBM Software Support | Enterprise software support | watsonx.ai; watsonx Orchestrate – Support-case answer generation and diagnosis | About ten minutes saved per eligible case during four months; projected annual savings of roughly 17,000 hours, with deployment to 2,800 support engineers. | 95 / Strong | Source |
| TBS-P026 | Ma’aden | Mining | Microsoft 365 Copilot; Copilot Studio; Azure OpenAI Service – Employee productivity, policy search, documents, presentations and meetings | 36,600+ Copilot interactions and up to 2,200 employee hours saved in one month during the first phase. | 80 / Strong | Source |
| TBS-P027 | Localiza&Co | Mobility / vehicle services | Microsoft 365 Copilot – Employee productivity, feedback analysis, meetings, marketing and accessibility | Average saving of 8.3 hours per employee monthly; 19 hours for the most engaged users; average reported productivity increase of 4.5%. | 90 / Strong | Source |
| TBS-P028 | Tüpraş | Energy / refining | Microsoft 365 Copilot – Employee productivity, meetings, documents and data analysis | The company estimates more than one hour saved daily per employee and 5,000 hours saved monthly across the workforce. | 70 / Moderate | Source |
| TBS-P029 | XP Inc. | Financial services | Microsoft 365 Copilot – Meetings, audit, governance, sales and accessibility | More than 9,000 hours saved; audit-team efficiency increased 30%; governance work saved 7% of time. | 92 / Strong | Source |
| TBS-P030 | Capita | Professional and business services | Microsoft 365 Copilot; Copilot Agents – Employee productivity, accessibility, knowledge search and no-code agents | 9,000 employee hours saved in one month after deployment of 3,000 licenses; employees built 169+ agents. | 77 / Moderate | Source |
| TBS-P031 | Bynder | Digital asset management software | Amazon Bedrock; Amazon Titan Multimodal Embeddings – Visual and natural-language asset search | One customer reported a 75% reduction in asset-search time and approximately 50% more options returned per search. | 63 / Moderate | Source |
| TBS-P032 | Saks | Luxury retail | Amazon Bedrock; Amazon Transcribe; Amazon Connect – Contact-centre call summarization and agent analytics | After-call work fell by 15 seconds per interaction across operations handling millions of customer contacts. | 77 / Moderate | Source |
| TBS-P033 | Dovetail | Customer-insights software | Amazon Bedrock – AI-assisted analysis of customer research | Customers reportedly save 10 hours weekly on data analysis and improve productivity by 80%; Dovetail can prototype in under a day and release features in two weeks. | 75 / Moderate | Source |
| TBS-P034 | KOHO | Financial technology | Amazon Bedrock and internal Kortex AI – Software delivery, suspicious-transaction reporting and security operations | 66% more frequent deployments, 3× faster suspicious-transaction report processing, 140% higher pull-request throughput and 50% faster security-event resolution. | 67 / Moderate | Source |
| TBS-P035 | Haast | Compliance software | Amazon Bedrock – Automated marketing and content-compliance review | Customers reportedly save 500+ hours of manual review, improve productivity by 80%+, launch campaigns 3× faster and analyse 4× longer videos. | 65 / Moderate | Source |
| TBS-P036 | GEMS | Mining technology | Gemini; Gemini Enterprise Agent Platform; Google Cloud – Pit-to-port operational intelligence and executive decision support | Executive decision-making accelerated by over 90%; data retrieval fell from two days to under one hour across 50+ applications and 4,000+ users. | 82 / Strong | Source |
| TBS-P037 | RELEX Solutions | Retail and supply-chain software | Vertex AI; Gemini – Promotion optimization and retail campaign-management agent | Campaign management reportedly fell from 4+ hours to under two minutes; retailers save tens of millions of euros annually and gain 2–3% net profit from optimized pricing. | 75 / Moderate | Source |
| TBS-P038 | Decidr | Enterprise data and agent software | Gemini; Google Workspace; Google Cloud – Data ingestion, integration and operational intelligence | Usable intelligence reportedly fell from months to hours; integration time declined up to 65%, onboarding complexity 50% and infrastructure overhead 45%; real-time data processing increased 10×. | 60 / Moderate | Source |
| TBS-P039 | Callers | Conversational AI software | Gemini; Vertex AI; Google Cloud – Voice and chat agents for lead conversion and customer engagement | Platform capacity scaled 30× to 250,000+ daily interactions; clients reported 50% higher conversion, 32% more closed deals in one month, over US$500,000 in revived leads and up to 80% lower campaign costs. | 85 / Strong | Source |
| TBS-P040 | Gelato | Global ecommerce production software | Gemini 1.5 Pro; Vertex AI – Support-ticket triage and machine-learning product development | 120 hours of manual work saved weekly; ticket-triage accuracy increased from 60% to 90%; model deployment fell from weeks to two days. | 77 / Moderate | Source |
| TBS-P041 | Quantium | Data analytics and professional services | Claude Enterprise / Claude Platform – Coding, proposals, training and leadership coaching | 89% of 1,200+ staff use AI daily; proposal work was cut by up to 90%; leadership-program development fell from 64 to 32 days. | 82 / Strong | Source |
| TBS-P042 | BlueFlame AI | Investment technology | Claude Platform – Financial-document and deal-room analysis | Document analysis reportedly fell from 4+ hours to minutes; the platform supports about 30 client queries daily and can compare hundreds of companies at once. | 80 / Strong | Source |
| TBS-P043 | Intercom | Customer-service software | Claude-powered Fin AI Agent – Automated customer support resolution | Fin reportedly averages a 51% resolution rate. Synthesia resolved 6,000+ conversations and saved 1,300+ hours in six months; Fundrise automated 50%+ of support volume with 95% accuracy in three months. | 90 / Strong | Source |
| TBS-P044 | Factory | Software-engineering automation | Claude-powered Factory Droids – Autonomous software development and code maintenance | Factory reports 550,000 engineering hours saved across customers, 20% shorter development cycles and a 3× reduction in code churn. | 65 / Moderate | Source |
| TBS-P045 | AppFolio | Property-management software | Claude in Amazon Bedrock; Realm-X – AI-assisted property-management communication | Property managers reportedly save 11 hours weekly; more than 500,000 AI-suggested responses were processed; response time fell by 26 seconds and adoption increased 3×. | 77 / Moderate | Source |
| TBS-P046 | ENEOS Materials | Advanced materials manufacturing | ChatGPT Enterprise – Research, production technology, HR analysis and company knowledge work | 80% of pilot employees reported significant workflow improvement; HR data aggregation and analysis time fell 90%; investigations reportedly fell from months to minutes. | 83 / Strong | Source |
| TBS-P047 | Paf | Gaming and technology | ChatGPT Enterprise; custom GPTs – Software development and companywide knowledge work | 70% of employees actively use ChatGPT Enterprise; 100 developers created 85 custom GPTs; the company estimates output equivalent to 12 full-time employees. | 80 / Strong | Source |
| TBS-P048 | Zenken | Marketing and human-resources services | ChatGPT Enterprise; GPT-5; custom GPTs – Sales preparation, research, translation, content and proposals | 90%+ weekly active usage; 30–50% average task-time savings; 5–15 hours released per employee monthly and roughly ¥50 million in annual outsourcing savings. | 82 / Strong | Source |
| TBS-P049 | Holiday Extras | Travel services | ChatGPT Enterprise; OpenAI API – Software engineering and companywide productivity | More than 500 hours saved weekly across the company; code-debugging time fell 75%; the company equates the saving to about US$500,000 annually. | 90 / Strong | Source |
| TBS-P050 | Morgan Stanley Wealth Management | Financial services | GPT-4; AI @ Morgan Stanley Assistant and Debrief – Advisor knowledge retrieval and meeting summarization | More than 98% of advisor teams use the assistant; document access rose from 20% to 80%; the searchable corpus expanded from 7,000 questions to 100,000 documents and client follow-ups fell from days to hours. | 90 / Strong | Source |
| TBS-P051 | AddAI | Conversational AI services | watsonx.ai; watsonx Assistant; Watson Discovery – Banking customer-service assistant | The pilot achieved 85% accuracy against predefined correct answers and reduced unanswered customer queries by 50%; implementation time is expected to fall up to 30%. | 73 / Moderate | Source |
| TBS-P052 | System Research | Systems integration | watsonx.ai; Watson Discovery – Internal natural-language document search | A three-week UI pilot produced a 50% reduction in the time employees needed to find information. | 65 / Moderate | Source |
| TBS-P053 | IBM Finance | Corporate finance | watsonx.ai; watsonx Orchestrate; RPA – Journal-entry processing, reconciliation and anomaly detection | During the initial three months, cycle time was projected to fall by more than 90% and annual cost savings were projected at about US$600,000. | 73 / Moderate | Source |
| TBS-P054 | CXReview | Customer-experience analytics | watsonx.ai – Call-transcript summarization and disposition automation | The first pilot phase is estimated to save about 23 agent hours per day by removing manual call-disposition comments. | 58 / Limited | Source |
| TBS-P055 | Enztec | Medical-device research | watsonx.ai; Watson Discovery – Medical-literature review and research summarization | Research productivity reportedly increased by more than 500%, reducing research time from days to hours. | 75 / Moderate | Source |
| TBS-P056 | Klarna | Financial technology / payments | OpenAI-powered customer-service assistant and internal AI tools – Customer service and companywide operating productivity | Klarna reported that 90% of employees had integrated AI into daily workflows, operating expenses fell 11%, and its customer-service assistant was associated with US$40 million in annualised savings. | 90 / Strong | Source |
| TBS-P057 | TELUS | Telecommunications | Fuel iX and secure internal generative-AI tools – Companywide employee productivity, translation, knowledge search and workflow assistance | TELUS reported more than 57,000 employees using 31 secure AI tools, an average of 40 minutes saved per use, and more than 500,000 employee hours reclaimed. | 80 / Strong | Source |
| TBS-P058 | TELUS | Telecommunications | Fuel iX customer-support tool on Azure OpenAI Service – Customer self-service and support-query resolution | TELUS reported more than 50,000 customer queries and a 28% increase in customers successfully finding information compared with conventional website search. | 85 / Strong | Source |
| TBS-P059 | PwC | Professional services | Microsoft 365 Copilot and firmwide AI tools – Companywide knowledge work, documents, communication and professional-services workflows | PwC reported deployment to more than 285,000 users in over 100 countries, more than 20 million Copilot actions in a month, and more than one million hours of capacity created. | 75 / Moderate | Source |
| TBS-P060 | KPMG UK | Professional services | Internal generative-AI tools – Professional-services productivity and knowledge work | KPMG UK reported a 15% productivity boost over the preceding year and 1.3 million employee prompts in one month. | 73 / Moderate | Source |
| TBS-P061 | ServiceNow | Enterprise software | Now Assist used internally through Now on Now – Employee self-service, case resolution, software development and knowledge workflows | ServiceNow reported US$10 million in annualised tangible benefit in the first 120 days, 50 FTEs of annualised productivity, 14% self-service deflection, and roughly 80% less time producing resolution notes. | 90 / Strong | Source |
| TBS-P062 | Octopus Energy | Energy retail | Arlo generative-AI customer-email assistant – Drafting customer-service email responses | In a three-month trial handling about 8,000 emails each week, AI-assisted replies achieved 76% customer satisfaction compared with 72% for comparable human-written responses. | 85 / Strong | Source |
| TBS-P063 | Accenture | Professional services | Microsoft 365 Copilot – Internal IT and employee knowledge work | Accenture reported that some Copilot users saved up to three hours per day and experienced improved work quality. | 63 / Moderate | Source |
| TBS-P064 | Repsol | Energy | Microsoft 365 Copilot – Information search, summarization, meetings and knowledge work | A four-month Repsol study reported an average saving of 121 minutes per employee each week; two out of three participants said they would not return to working without the tool. | 83 / Strong | Source |
| TBS-P065 | Deloitte Tohmatsu Group | Professional services | Internal generative-AI environment – Companywide professional-services and knowledge workflows | Deloitte Tohmatsu reported approximately 12,000 staff using its internal generative-AI environment and about 100,000 employee hours saved each month. | 77 / Moderate | Source |
| TBS-P066 | Deloitte US | Professional services | Sidekick internal generative-AI platform – Firmwide knowledge work and reusable employee-built AI skills | Deloitte reported about two hours saved per user each week and more than 3,000 employee-created reusable skills. | 77 / Moderate | Source |
| TBS-P067 | Nilfisk | Industrial manufacturing | Microsoft 365 Copilot – Cross-functional knowledge work across 12 business areas | KPMG reported a 22% productivity improvement and 21% quality improvement across 192 identified use cases; marketing and product teams recorded improvements of about 35%. | 83 / Strong | Source |
| TBS-P068 | Telia | Telecommunications | Generative-AI data catalogue and enrichment workflow – Enterprise data tagging, descriptions and metadata management | Accenture reported that the system updated tags and descriptions for 87,000 data fields, saved 10,000 data-engineer hours and avoided about €700,000 in data-entry costs. | 85 / Strong | Source |
| TBS-P069 | Navantia | Shipbuilding and defence manufacturing | Microsoft 365 Copilot – Employee collaboration, documents, meetings and information access | Capgemini reported 91.7% licence activation, an average saving of 30 minutes per user per day and user satisfaction of 8.3 out of 10. | 85 / Strong | Source |
| TBS-P070 | Siemens industrial customers | Industrial manufacturing | Siemens Industrial Copilot – Industrial engineering, automation-code generation and panel visualisation | Siemens reported more than 100 customer companies using the Industrial Copilot, potential access for 120,000 engineers, panel visualisations produced in 30 seconds and generated code requiring about 20% adaptation. | 85 / Strong | Source |
| TBS-P071 | Procter & Gamble professionals | Consumer goods / product innovation | GPT-4-based generative-AI assistance – Product-innovation problem solving by individuals and teams | A randomised experiment with 776 P&G professionals found that individuals using AI achieved performance comparable with conventional two-person teams and that AI reduced functional knowledge silos between commercial and R&D participants. | 90 / Strong | Source |
| TBS-P072 | Microsoft internal developer-support service | Technology | GPT-3.5 and GPT-4 support bots – Internal developer information retrieval and support escalation | Across 3,296 sessions, GPT-based bots reduced escalation by 9.2 percentage points, a 53.8% relative reduction from the classical bot’s 17.1% baseline. | 85 / Strong | Source |
| TBS-P073 | Ant Group programmer teams | Financial technology | CodeFuse generative-AI coding assistant – Software development | The study reported a 55% increase in lines of code produced among AI users, with statistically significant gains concentrated among entry-level and junior programmers. | 85 / Strong | Source |
| TBS-P074 | Google software engineers | Technology | Internal Google AI coding features – Complex enterprise-grade software-development task | A trial with 96 full-time Google engineers estimated that AI shortened time on a complex enterprise-grade development task by about 21%, although the confidence interval was wide. | 85 / Strong | Source |
| TBS-P075 | Cross-industry knowledge workers | Multiple industries | Generative-AI tool integrated into email, documents and meetings – Everyday enterprise information work | In a 6,000-worker trial, active users spent three fewer hours, or 25% less time, on email each week; the intent-to-treat estimate was 1.4 hours, documents were completed moderately faster and meeting time did not change significantly. | 90 / Strong | Source |
| TBS-P076 | Microsoft, Accenture and an anonymous Fortune 100 company | Technology / professional services / diversified enterprise | AI coding assistant – Ordinary-course software development | Across 4,867 developers, combined analysis found a 26.08% increase in completed tasks, with higher adoption and larger gains among less experienced developers. | 90 / Strong | Source |
| TBS-P077 | Fortune 500 retailer employees | Retail | Enterprise generative-AI tool – Document production and structured human-AI collaboration | Among 388 employees, a mandatory paired-use protocol was associated with lower document quality and substantially lower production, while cognitive reframing improved document quality at the top of the distribution. | 70 / Moderate | Source |
| TBS-P078 | Penn Medicine clinicians | Healthcare | Ambient AI clinical documentation tool – Clinical note generation and after-hours documentation | Among 46 clinicians, time in notes fell from 10.3 to 8.2 minutes per appointment, same-day closure rose from 66.2% to 72.4%, and after-hours documentation fell from 50.6 to 35.4 minutes. | 95 / Strong | Source |
| TBS-P079 | Clinicians across six US health systems | Healthcare | Ambient AI clinical scribe – Clinical documentation and clinician workload | Across 263 clinicians, burnout declined from 51.9% before implementation to 38.8% after 30 days, alongside reductions in cognitive load and after-hours documentation. | 95 / Strong | Source |
| TBS-P080 | Northern and Central California clinicians | Healthcare | Ambient AI documentation technology – Clinical documentation, cognitive load and burnout | Among 100 participating clinicians, burnout declined from 42.1% to 35.1%, although the change was not statistically significant; electronic-record metrics and surveys showed improvements in documentation time, cognitive load and satisfaction. | 95 / Strong | Source |
| TBS-P081 | JPMorganChase | Financial services | LLM Suite – Secure companywide generative-AI access for drafting, research and idea generation | JPMorganChase reported that LLM Suite expanded to more than 200,000 onboarded employees within eight months, with 65,000 active users in the Corporate & Investment Bank. | 65 / Moderate | Source |
| TBS-P082 | JPMorganChase | Financial services | AI transaction-screening systems – Transaction screening and manual review reduction | JPMorganChase reported that AI enabled transaction-screening teams to review more than twice the previous volume while cutting manual operator checks by half. | 83 / Strong | Source |
| TBS-P083 | JPMorgan Asset & Wealth Management | Financial services | Connect Coach – Advisor content retrieval, call summaries and client preparation | The 2024 annual report states that Connect Coach was approximately 95% quicker at locating relevant content for client conversations. | 80 / Strong | Source |
| TBS-P084 | JPMorgan Asset & Wealth Management | Financial services | Sales Assist – Real-time sales support, call transcription and client information | JPMorganChase reported approximately 20% higher gross sales supported by Sales Assist year over year. | 78 / Moderate | Source |
| TBS-P085 | Vodafone | Telecommunications | SuperTOBi generative-AI virtual assistant – Customer-service resolution and virtual assistance | Vodafone reported about 60 million virtual-assistant conversations monthly, a 70% end-to-end resolution rate and an eight-percentage-point NPS improvement over the previous AI agent. | 95 / Strong | Source |
| TBS-P086 | Vodafone | Telecommunications | Autonomous generative-AI procurement portal – Tender sourcing and procurement workflows | Vodafone reported that its generative-AI sourcing platform was used in more than 90% of tenders, served more than 10,000 suppliers and reduced sourcing time by more than 30%. | 92 / Strong | Source |
| TBS-P087 | Vodafone | Telecommunications | Generative-AI software-development tools – Software coding and development lifecycle assistance | Vodafone reported more than 5,000 active developers, over 35% AI-code acceptance and a 12% productivity increase across the development lifecycle. | 90 / Strong | Source |
| TBS-P088 | Vodafone | Telecommunications | Generative-AI HR digital agent – Personalised multilingual employee support | Vodafone reported that its HR digital agent served 75,000 employees and resolved 86% of requests. | 87 / Strong | Source |
| TBS-P089 | UK Civil Service cross-government cohort | Government / public administration | Microsoft 365 Copilot – Documents, email, meetings, search and routine knowledge work | A three-month experiment involving 20,000 employees across 12 organizations found self-reported average savings of 26 minutes per day; 82% did not want to return to pre-Copilot working conditions. | 90 / Strong | Source |
| TBS-P090 | UK Department for Work and Pensions | Government / public administration | Microsoft 365 Copilot – Search, email, documents, meeting summaries and administrative work | Among 3,549 licence holders, econometric analysis estimated an average saving of 19 minutes per day; 65% reported greater fulfilment, with a 0.56-point job-satisfaction increase versus non-users. | 95 / Strong | Source |
| TBS-P091 | HM Revenue & Customs and Valuation Office Agency | Government / tax administration | Microsoft 365 Copilot – Documents, meetings, search, coding and spreadsheet work | The Phase III trial reported approximately 60 minutes saved per week after conservative adjustments; analysts estimated that up to 50,000 licenses could generate about £50 million in annual net productivity capacity. | 95 / Strong | Source |
| TBS-P092 | NHS organizations | Healthcare | Microsoft 365 Copilot – Healthcare administration and knowledge work | A pilot involving more than 30,000 workers across 90 NHS organizations reported average savings of at least 43 minutes per person per day and projected up to 400,000 hours saved monthly under full rollout. | 80 / Strong | Source |
| TBS-P093 | UK public-sector software developers | Government / technology | GitHub Copilot and other AI coding assistants – Software coding, search and technical problem solving | More than 1,000 public-sector technology specialists reported saving almost one hour per day, equivalent to 28 working days annually; 15.8% of suggested code lines were accepted. | 90 / Strong | Source |
| TBS-P094 | Accenture developers | Professional services / technology | GitHub Copilot – Enterprise software development | GitHub reported an 8.69% increase in pull requests, a 15% increase in pull-request merge rate and an 84% increase in successful builds among Accenture developers using Copilot. | 80 / Strong | Source |
| TBS-P095 | Amazon | Technology / ecommerce | Amazon Q Developer code transformation – Modernising internal Java production applications | Amazon reported migrating tens of thousands of production applications, saving more than 4,500 developer-years and generating US$260 million in annual performance-related cost savings. | 100 / Strong | Source |
| TBS-P096 | Amazon | Technology / ecommerce | Amazon Q Business internal knowledge assistant – Developer technical-question resolution and internal knowledge retrieval | Amazon reported more than one million internal developer questions resolved in 2024 and more than 450,000 hours of technical-investigation time saved, with answer time falling from hours to seconds. | 100 / Strong | Source |
| TBS-P097 | Salesforce | Enterprise software | Agentforce in Slack – Employee questions and actions inside Slack | Salesforce reported 86% employee adoption after six months and said Agentforce in Slack was on track to save more than 500,000 employee hours annually. | 85 / Strong | Source |
| TBS-P098 | Salesforce | Enterprise software | Agentforce HR Service and Employee Portal – Employee HR self-service and case support | Salesforce reported nearly 10 million employee-portal searches and a 96% self-service resolution rate without HR-team intervention. | 87 / Strong | Source |
| TBS-P099 | Salesforce | Enterprise software | Agentforce Service Agent – Customer-support conversations and human-capacity reallocation | Salesforce reported that Agentforce handled more than two million customer-support conversations and now manages the bulk of daily conversations on its help site. | 85 / Strong | Source |
| TBS-P100 | Salesforce | Enterprise software | Agentforce engineering and incident-response agents – Code maintenance, threat monitoring and incident remediation | Salesforce reported a 30% improvement in engineering cycle time, with agents detecting 91% of incidents within eight minutes and automatically remediating 87% within 20 minutes. | 87 / Strong | Source |
| TBS-P101 | SAP service consultants | Enterprise software | Joule – Service-consultant knowledge retrieval and customer support | SAP reported that up to 4,000 internal service consultants using Joule could save as much as two hours per day. | 75 / Moderate | Source |
| TBS-P102 | McKinsey & Company | Professional services | Lilli – Knowledge search, synthesis, presentations, proposals and research | McKinsey reported 92% workforce adoption, 74% regular use, more than 30% time saved on information gathering and synthesis, nearly 19 million prompts, and two to three million cumulative hours saved. | 90 / Strong | Source |
| TBS-P103 | Moderna | Biotechnology / pharmaceuticals | ChatGPT Enterprise and internal GPTs – Legal, research, manufacturing, commercial and clinical-support workflows | Moderna reported more than 80% internal adoption and deployment of more than 750 purpose-built GPTs across business functions within a few months. | 60 / Moderate | Source |
| TBS-P104 | Bayer Crop Science | Agriculture / life sciences | E.L.Y. agentic generative-AI assistant – Agronomic and product-question support for frontline staff | Bayer reported a 60% improvement in response times and employee time savings of up to four hours per week based on internal benchmarks. | 83 / Strong | Source |
| TBS-P105 | Bosch manufacturing plant | Industrial manufacturing | Generative synthetic-image system for visual inspection – Training automated optical-inspection models for electric-motor production | Bosch reported generating about 15,000 synthetic defect images, shortening the project by an expected six months and projecting annual productivity gains in the six-figure euro range. | 90 / Strong | Source |
| TBS-P106 | Walmart | Retail | AI-directed task-management workflow – Shift planning and frontline task prioritisation | Walmart reported that early use reduced team-lead shift-planning time from 90 minutes to 30 minutes. | 78 / Moderate | Source |
| TBS-P107 | Walmart | Retail | RFID and augmented-reality product-search system – Locating merchandise in stores and backrooms | Walmart reported that the RFID-and-AR workflow made product searches up to 75% faster. | 73 / Moderate | Source |
| TBS-P108 | Target | Retail | Store Companion generative-AI chatbot – Store-process questions, coaching and operations support | Target announced a planned rollout to hundreds of thousands of team members across nearly 2,000 stores after piloting the tool in about 400 stores, but did not publish a numerical business outcome. | 50 / Limited | Source |
| TBS-P109 | BMW Group | Automotive manufacturing | Virtual Factory digital twins with generative and agentic AI – Production planning, collision checks, layout and logistics simulation | BMW projected that its Virtual Factory could reduce production-planning costs by up to 30%, with digital twins covering more than 30 production sites. | 95 / Strong | Source |
| TBS-P110 | Siemens industrial customers | Industrial automation | Eigen Engineering Agent – End-to-end automation engineering and project onboarding | Siemens reported two-to-five-times faster engineering workflows, up to 80% higher solution quality and 50% greater engineering efficiency after pilots with more than 100 customers in 19 countries. | 85 / Strong | Source |
| TBS-P111 | Mercedes-Benz | Automotive manufacturing | AI-controlled process engineering – Monitoring and controlling automotive paint-shop processes | Mercedes-Benz reported a 20% reduction in energy consumption after replacing conventional controls with AI-controlled process engineering in top-coat booths. | 73 / Moderate | Source |
| TBS-P112 | Nestlé | Consumer goods / food and beverage | Agentic-AI virtual sales assistant – Automating routine sales tasks in pilot markets | Nestlé reported that the assistant automated up to 40% of routine tasks and saved salespeople 20% to 35% of the time spent on those tasks in pilot markets. | 73 / Moderate | Source |
| TBS-P113 | Nestlé | Consumer goods / food and beverage | AI-powered marketing content adaptation – Adapting marketing assets across brands, formats and markets | Nestlé reported up to a 75% reduction in content-adaptation cost and more than a 70% increase in delivery speed. | 83 / Strong | Source |
| TBS-P114 | Nestlé | Consumer goods / food and beverage | Generative-AI digital-twin content service – Creating product imagery for ecommerce and digital media | Nestlé reported that the service reduced the time and cost of scaling digital twins by more than 70%. | 95 / Strong | Source |
| TBS-P115 | A.P. Moller – Maersk | Logistics | AI Agent Assist – Generating customer-service email recommendations | Maersk reported that the pilot generated pre-written recommendations for more than 80% of queries within an operation handling 28 million emails annually, with an average manual resolution time of five minutes. | 80 / Strong | Source |
| TBS-P116 | BT Group | Telecommunications | Amazon CodeWhisperer – Software-code generation and developer assistance | BT reported more than 100,000 generated lines of code in four months, automating about 12% of repetitive work for initial volunteers, with a 37% suggestion-acceptance rate. | 95 / Strong | Source |
| TBS-P117 | BT Group | Telecommunications | ServiceNow Now Assist – Writing case summaries and reviewing complex customer-service notes | BT reported that Now Assist cut both case-summary writing time and complex-note review time by 55%. | 58 / Limited | Source |
| TBS-P118 | DBS Bank | Banking | DBS Joy generative-AI chatbot – Corporate and SME customer support | DBS reported more than 20,000 unique users, about 15,000 monthly chat sessions and a 23% improvement in customer satisfaction over six months. | 95 / Strong | Source |
| TBS-P119 | DBS Bank | Banking | Generative-AI trade-condition processing – Processing trade-finance conditions and documentation | DBS reported a 60% reduction in trade-condition processing time. | 78 / Moderate | Source |
| TBS-P120 | DBS Bank | Banking | Generative-AI KYC name-screening workflow – Know Your Customer screening | DBS reported approximately 70% efficiency gains in name screening. | 73 / Moderate | Source |
| TBS-P121 | Providence St Joseph Health clinicians | Healthcare | Ambient AI clinical documentation tools – Clinical documentation burden, productivity and efficiency | Across 1,547 clinicians and 16,149 observations, simple pre-post analysis found time in notes per appointment fell from 7.1 to 6.1 minutes; interrupted-time-series analysis also identified sustained documentation changes. | 100 / Strong | Source |
| TBS-P122 | UCSF Health physicians | Healthcare | Ambient AI clinical scribes – Documentation and physician financial productivity | Among 1,565 physicians and 1,202,734 encounters, adopters gained 1.81 relative-value units per week and 0.80 encounters per week, equivalent to an estimated US$3,044 annually per physician. | 100 / Strong | Source |
| TBS-P123 | UC San Diego Health sepsis-care teams | Healthcare | LLM-assisted sepsis-record abstraction and feedback – Real-time sepsis quality-measure feedback | Among 66 physicians treating 301 patients, SEP-1 compliance increased from 70.1% in the control group to 82.9% with the intervention, a 13-point absolute improvement. | 90 / Strong | Source |
| TBS-P124 | UCLA Health outpatient physicians | Healthcare | Microsoft DAX Copilot and Nabla ambient scribes – Clinical documentation time and physician wellbeing | Among 238 physicians, Nabla reduced time in notes by 9.5% versus usual care, while DAX showed no significant change; both tools produced possible improvements in burnout-related measures. | 90 / Strong | Source |
| TBS-P125 | University of Wisconsin health practitioners | Healthcare | Ambient AI clinical documentation tool – Documentation burden, wellbeing and billing quality | Across 66 practitioners and 71,487 notes, ambient AI reduced work exhaustion by 0.44 points and time spent on notes by 0.36 hours per day without reducing documentation quality. | 90 / Strong | Source |
| TBS-P126 | Fortune 500 customer-support operation | Business-process software / customer service | GPT-3-based conversational assistant – Real-time recommendations for customer-support agents | Across 5,172 agents and more than three million chats, AI access increased issues resolved per hour by 15% on average; lower-skilled and less-experienced agents gained substantially more. | 95 / Strong | Source |
| TBS-P127 | AXA | Insurance | AXA Secure GPT – Secure companywide access to generative AI | AXA stated that Secure GPT was available to 150,000 collaborators, but the cited source did not publish a numerical business outcome. | 65 / Moderate | Source |
| TBS-P128 | Sanofi | Pharmaceuticals | Concierge internal generative-AI platform – Enterprise knowledge work and internal AI assistance | Sanofi stated that Concierge had been used by more than 50,000 employees over approximately 18 months, but published no numerical productivity or financial outcome in the source. | 65 / Moderate | Source |
| TBS-P129 | DHL Group | Logistics | Autonomous agentic-AI communication systems – Phone, email and messaging automation in logistics operations | DHL reported that initial deployments handled hundreds of thousands of emails and millions of call minutes annually, with 11 million phone-call minutes identified for automation. | 70 / Moderate | Source |
| TBS-P130 | DBS Bank | Banking | CodeBuddy generative and agentic AI assistant – Software coding for data professionals | DBS reported time savings of up to 20% on selected coding tasks. | 73 / Moderate | Source |
| TBS-P131 | Commonwealth of Pennsylvania | Government / public administration | ChatGPT Enterprise – Writing, research, summarization, IT support and administrative work | A year-long pilot with 175 employees across 14 agencies reported average self-assessed savings of 95 minutes on days employees used ChatGPT; more than 85% described the experience positively. | 90 / Strong | Source |
| TBS-P132 | Australian Public Service | Government / public administration | Microsoft 365 Copilot – Summarisation, drafting, rewriting and information search | A six-month trial with 7,600 users across more than 60 agencies found 69% believed Copilot improved task speed and 61% believed it improved work quality; some tasks produced perceived savings of about one hour, while up to 7% said the tool added time. | 95 / Strong | Source |
| TBS-P133 | CSIRO | Scientific research | Microsoft 365 Copilot – Meetings, email drafting, content retrieval and research administration | A 300-licence trial found productivity and efficiency improvements in structured tasks, but broader effectiveness was constrained by integration gaps, advanced-function limitations and increased post-trial concerns about privacy and ethics. | 68 / Moderate | Source |
| TBS-P134 | Singapore public-service prototype users | Government / public administration | Insight survey-analysis agent – Classifying and analysing citizen-feedback and survey data | Testing classified 200 ScamShield issues in about one minute and completed trend analysis in about three minutes, reducing a workflow described as four days to four minutes, or roughly 99%. | 85 / Strong | Source |
| TBS-P135 | GovTech Singapore public-sector developers | Government / software development | Generative-AI coding assistants – Public-sector software development | The study reported coding or task-speed improvements of 21% to 28% and 95% developer satisfaction, with particularly strong gains for junior developers. | 78 / Moderate | Source |
| TBS-P136 | Enterprise IT administrators | Information technology administration | Microsoft Security Copilot – Sign-in troubleshooting, device-policy management and device troubleshooting | Across three IT-administration scenarios, Security Copilot users improved overall accuracy by 34.53% and reduced task-completion time by 29.79%; free-response tasks showed a 61.14% time reduction. | 80 / Strong | Source |
| TBS-P137 | McDonald’s | Restaurants / food service | IBM automated drive-through order taker – Voice-based drive-through ordering | McDonald’s ended a roughly two-and-a-half-year AI order-taking test at more than 100 restaurants and removed the system, while saying it would continue evaluating other voice-ordering options. | 60 / Moderate | Source |
| TBS-P138 | DPD | Logistics / parcel delivery | AI customer-service chatbot – Parcel-status and customer-service assistance | DPD disabled the AI component of its customer-service chatbot after a system update allowed it to swear, criticise the company and fail to resolve a customer’s parcel query. | 48 / Limited | Source |
| TBS-P139 | Air Canada | Airlines | Customer-service chatbot – Fare-policy and bereavement-discount information | A tribunal held Air Canada responsible for misleading information supplied by its chatbot and ordered the airline to pay C$812.02 in damages, interest and fees. | 95 / Strong | Source |
| TBS-P140 | Cando Rail & Terminals | Rail transportation | Internal safety chatbot – Summarising rail-safety rules and training material | Cando paused an internal chatbot after it inconsistently summarized a roughly 100-page safety rulebook, sometimes forgetting, misinterpreting or inventing rules; the company had spent about US$300,000 developing AI products. | 70 / Moderate | Source |
| TBS-P141 | CellarTracker | Consumer software / wine data | OpenAI-based AI sommelier – Personalised wine recommendations | CellarTracker required six weeks of prompt and guardrail tuning before launch because its AI sommelier was overly agreeable and would not reliably tell users they were unlikely to enjoy a wine. | 68 / Moderate | Source |
| TBS-P142 | Verizon | Telecommunications | AI-assisted customer-service routing and triage – Screening calls, customer context and routing between self-service and human agents | Verizon retained roughly 2,000 frontline service agents and reported that about 40% of consumers still preferred human access; AI remained focused on triage and routine tasks rather than holistic customer conversations. | 80 / Strong | Source |
| TBS-P144 | New York City MyCity | Government / small-business services | MyCity business-information chatbot – Answering business questions about city laws, permits and regulations | Independent testing found that the chatbot repeatedly supplied inaccurate or potentially illegal advice on housing, worker rights and business rules; the city added stronger disclaimers and later planned to end the beta. | 63 / Moderate | Source |
| TBS-P145 | Lloyds Banking Group | Banking | Portfolio of generative and agentic AI systems – Customer operations, coding, risk, HR and banking workflows | Lloyds reported that generative AI delivered approximately £50 million of value in 2025 across more than 50 deployed use cases, with more than £100 million of additional value expected in 2026. | 90 / Strong | Source |
| TBS-P146 | Lloyds Banking Group | Banking | Athena generative-AI knowledge hub – Customer-service knowledge retrieval | Athena reduced average information-search time from 59 seconds to about 20 seconds, a 66% reduction, and was expected to save 4,000 hours for telephone-banking teams; 21,000 colleagues made 2.1 million searches. | 100 / Strong | Source |
| TBS-P147 | Lloyds Banking Group | Banking | Microsoft 365 Copilot – Documents, meetings and administrative knowledge work | Nearly 30,000 licenses were deployed, 93% of licensed employees were active, and an internal survey reported average savings of 46 minutes per day. | 95 / Strong | Source |
| TBS-P148 | Lloyds Banking Group | Banking | Dynamic Risk Engine – Fraud and financial-crime risk identification | Lloyds reported that its machine-learning Dynamic Risk Engine identified £30 million of previously unknown high-value fraud risk. | 75 / Moderate | Source |
| TBS-P149 | HSBC | Banking | AI payment-fraud risk scoring – Payment-fraud detection and false-positive reduction | HSBC reported a 60% reduction in false positives after deploying AI risk scoring based on transaction, account, KYC and suspicious-activity data. | 78 / Moderate | Source |
| TBS-P150 | NatWest Group | Banking | Cora generative-AI digital assistant – Customer-service queries, fraud reporting and adviser escalation | NatWest said generative-AI functionality drove a 150% improvement in customer-satisfaction levels and reduced the need for a human adviser to complete requests; the earlier Cora service handled 10.8 million queries in 2023. | 85 / Strong | Source |
| TBS-P151 | Bank of America | Banking | Erica for Employees – Internal IT and employee-service support | More than 90% of employees used Erica for Employees, and Bank of America reported that the assistant reduced calls to the IT service desk by more than 50%. | 90 / Strong | Source |
| TBS-P152 | Bank of America | Banking | Generative-AI coding assistant – Code writing and optimization | Bank of America reported developer efficiency gains of more than 20%; later company data stated that approximately 19,000 developers used real-time coding assistance. | 90 / Strong | Source |
| TBS-P153 | Commonwealth Bank of Australia | Banking | Autonomous AI fraud-detection agent – Detecting emerging card fraud and creating or updating detection rules | CommBank reported that AI fraud technology helped reduce fraud losses by more than 20% in the first half of FY2026 versus the first half of FY2025, while the agent contributed to creating or updating three quarters of card-fraud rules. | 100 / Strong | Source |
| TBS-P154 | Manulife | Insurance | ChatMFC and enterprise AI portfolio – Companywide productivity, agent support, email drafting and coaching | Manulife reported 91 AI use cases in production, 121 more in development and 100% global workforce coverage for ChatMFC, but did not publish a portfolio-level productivity or financial return. | 65 / Moderate | Source |
| TBS-P155 | Rentokil Initial | Business services / pest control | Google Gemini and proprietary AI Portal – Employee productivity and creation of company-specific agents and chatbots | Rentokil reported that all colleagues worldwide had access to Gemini by the end of 2025 and that approximately 100 company-specific agents and chatbots were in development, but it disclosed no measured business outcome. | 60 / Moderate | Source |
| TBS-P156 | Mercer International | Biomaterials manufacturing | Google Workspace with Gemini and Google Vids – Companywide productivity and safety-training content production | Mercer projected US$3 million in annual productivity value after applying conservative filters to employee-reported savings, and reported a 75% reduction in safety-video production cost. | 90 / Strong | Source |
| TBS-P157 | Generali | Insurance | AI, automation and Amazon Bedrock-based generative-AI systems – Pricing, underwriting, claims, operations and customer assistance | Generali reported more than €200 million in operating-cost savings over three years across 16 flagship AI and automation use cases, while claim settlement fell from several days to one day or seconds for simple health claims. | 92 / Strong | Source |
| TBS-P158 | Forethought Technologies | Customer-service software | Amazon SageMaker inference for SupportGPT – Reducing generative-AI model inference and hosting cost | Forethought reduced model-serving costs by up to 66% using multi-model endpoints and about 80% for selected classifiers using serverless inference, while supporting more than 30 million customer interactions annually. | 90 / Strong | Source |
| TBS-P159 | Discover Financial Services | Financial services | Cloud data-science and generative-AI workbench – Large-scale model training, sentiment analysis and customer classification | Discover reduced processing of 20,000 records from 6.5–7 hours on CPUs to four minutes on multiple GPUs, reduced 30 million-record processing from days to hours and cut a 57,000-record sentiment analysis from hours to minutes. | 95 / Strong | Source |
| TBS-P160 | MIXI | Technology, gaming and digital services | ChatGPT Enterprise – Companywide knowledge work, coding and process improvement | MIXI estimated 17,600 work hours saved monthly, about 11 hours per active user, with 99% of surveyed users reporting improved productivity. | 95 / Strong | Source |
| TBS-P161 | Aberdeen City Council | Local government | Microsoft 365 Copilot – Meeting minutes, reports, policy guidance, translation and resident services | The council projected a 241% return on investment and approximately US$3 million in annual time and productivity value after deploying Copilot to 700 users. | 85 / Strong | Source |
| TBS-P162 | Quilter | Wealth management | Microsoft 365 Copilot – Post-call administration, meeting notes and document creation | Quilter estimated more than 13,000 hours saved monthly in post-call administration and reported reaching its technology breakeven point in one month. | 100 / Strong | Source |
| TBS-P163 | Reckitt | Consumer health and household products | Azure OpenAI and AI Marketing Insights Generator – Consumer insights, concept development, ad adaptation and campaign analysis | Reckitt reported at least a 60% marketing-efficiency boost, 30% faster ad adaptation, 60% faster concept development and up to 90% less time on routine marketing analysis. | 83 / Strong | Source |
| TBS-P164 | Auquan customers | Financial services | Auquan autonomous agents on Azure OpenAI – Financial document processing, research and report writing | Auquan reported eliminating up to 95% of manual effort, saving financial institutions more than 50,000 hours and reducing workflow costs by 50%. | 70 / Moderate | Source |
| TBS-P165 | Emccamp | Real estate and construction | Gemini, Imagen, Veo and agentic workflows – Landing pages, marketing production and land-acquisition research | Emccamp reduced landing-page cost from R$8,000 to R$50 and delivery from three weeks to four hours, estimated R$170,000–R$200,000 in marketing-production savings and cut land research from five weeks to minutes. | 85 / Strong | Source |
| TBS-P166 | Indiana Farmers Insurance leadership team | Insurance | Enterprise generative-AI tools and adoption programme – Leadership research, analysis, writing and decision support | A 16-week programme reported 218 hours saved per week across 44 leaders, equivalent to 5.4 FTEs and approximately US$1.05 million in annual time value; daily AI use reached 92%. | 90 / Strong | Source |
| TBS-P167 | Anonymous leading cross-border online retailer | Online retail | Seven consumer-facing generative-AI applications – Marketplace search, content and consumer shopping workflows | Across seven workflows involving millions of users and products, treatment effects on sales ranged from 0% to 16.3%; four positive applications produced an implied annual incremental value of about US$5 per consumer. | 95 / Strong | Source |
| TBS-P168 | Irish public-sector financial-regulation staff | Government / financial regulation | GPT-4o document and data-analysis applications – Document understanding and structured data analysis | Generative AI improved document-task quality by 17% and completion speed by 34%, but reduced data-analysis quality by 12% and did not significantly improve data-task completion time. | 85 / Strong | Source |
| TBS-P169 | Company S | Large enterprise / corporate finance | Generative AI, intelligent document processing and automation agent – End-to-end corporate expense and receipt processing | The system reduced paper-receipt expense-processing time by more than 80%, while also reducing errors and improving compliance through human-in-the-loop exception handling. | 78 / Moderate | Source |
| TBS-P170 | NRI and client organizations | Professional services / financial and industrial clients | Claude in Amazon Bedrock – Complex Japanese document review and internal testing | NRI reported a 50% reduction in review time for complex Japanese business documents, with up to 85% productivity improvement in testing and 40% in development processes. | 73 / Moderate | Source |
| TBS-P171 | Quillit customers | Market research / professional services | Claude Platform – Qualitative-research synthesis and report writing | Quillit reported an 80% reduction in report-writing time and citation accuracy of 89%–98% using Claude 3.5. | 68 / Moderate | Source |
| TBS-P172 | Inscribe customers | Financial risk technology | Claude-powered AI Risk Agents – Fraud review, document verification and risk analysis | Inscribe reported reducing fraud-review time from 30 minutes to 90 seconds, a 20-fold improvement, and increasing output 70-fold in one client example. | 75 / Moderate | Source |
| TBS-P173 | Graphite customers | Software development | Claude-powered AI code reviewer – Pull-request review, bug detection and fix suggestions | Graphite reported reducing pull-request feedback from one hour to 90 seconds, with 96% positive feedback, a 67% implementation rate for suggestions and support for hundreds of thousands of pull requests. | 90 / Strong | Source |
| TBS-P174 | Armanino | Accounting and consulting | Claude on Amazon Bedrock – Audit rejection notes, client communication and AI tool development | Armanino estimated a 65% reduction in manual writing time, a 60% reduction in follow-up clarifications and thousands of hours saved annually; its first AI tool moved from idea to production in two weeks. | 90 / Strong | Source |
| TBS-P175 | CodeRabbit customers | Software development | Claude-powered AI code review – Code review, quality assurance and security checks | CodeRabbit reported 86% faster code delivery, more than 60% fewer code-review issues, a 70% implementation rate for AI fixes and processing of millions of pull requests monthly. | 85 / Strong | Source |
| TBS-P176 | Norges Bank Investment Management | Sovereign asset management | Claude for financial analysis and data access – Portfolio analysis, earnings calls, news monitoring and voting | NBIM estimated approximately 20% productivity gains, equivalent to 213,000 hours, while using Claude to analyse earnings calls and monitor news across more than 9,000 portfolio companies. | 75 / Moderate | Source |
| TBS-P177 | Novo Nordisk | Pharmaceuticals | NovoScribe using Claude on Amazon Bedrock – Clinical-study reports, device protocols and patient materials | Novo Nordisk reported reducing clinical-study documentation from more than 10 weeks to 10 minutes and cutting review cycles by 50%; other materials that previously required months could be produced in under a minute. | 75 / Moderate | Source |
| TBS-P178 | Axon Draft One customers | Public safety technology | Draft One on Azure OpenAI Service – Drafting police incident reports from body-camera audio | Axon reported that Draft One cut report-writing time by about half for most officers and by up to 82% at some agencies. | 85 / Strong | Source |
| TBS-P179 | Mercado Libre | Online marketplace | Vertex AI Search – Product discovery across marketplace inventory | Mercado Libre deployed AI search across 150 million items in three pilot countries for a marketplace serving 100 million customers and reported millions of dollars in incremental revenue. | 80 / Strong | Source |
| TBS-P180 | Rakuten | Technology, ecommerce and financial services | Claude Code – Software development, feature delivery and autonomous refactoring | Rakuten reported reducing average feature-delivery time from 24 working days to five days, a 79% reduction; a complex refactoring task ran autonomously for seven hours and achieved 99.9% numerical accuracy. | 90 / Strong | Source |
| TBS-P181 | Commonwealth Bank of Australia | Banking | ChatIT using Microsoft AI and OpenAI – Internal IT support and automated issue remediation | ChatIT processed more than 2.3 million messages and executed over 12,000 automated IT fixes in six months, saving nearly 2,500 hours; each fix averaged two minutes versus a typical 17-minute service-desk call. | 100 / Strong | Source |
| TBS-P182 | Bunnings | Retail / home improvement | Internal staff chatbot – Store policy, product and operational question answering | Bunnings reported that its staff chatbot answered more than four million questions across 400 sites and 50,000 employees, saving staff an average of five minutes per day. | 80 / Strong | Source |
| TBS-P183 | Bunnings | Retail / home improvement | Buddy AI shopping assistant – Customer DIY guidance and product advice | Bunnings reported that 25,000 customers used Buddy in one week, but the disclosure did not publish conversion, revenue, resolution or satisfaction outcomes. | 60 / Moderate | Source |
| TBS-P184 | Westpac | Banking | Microsoft Copilot – Companywide knowledge work and employee experimentation | Westpac gave all 35,000 employees access to Copilot, but the source did not publish an aggregate time, cost, quality or revenue outcome. | 65 / Moderate | Source |
| TBS-P185 | Westpac | Banking | Real-time AI scam-call assistant – Detecting scam indicators and coaching frontline fraud specialists | Westpac said early pilot results indicated faster, more effective and more consistent scam support, but it did not publish a numerical detection, loss-prevention or handling-time outcome. | 45 / Limited | Source |
| TBS-P186 | ANZ | Banking | Microsoft 365 Copilot and GitHub Copilot – Email, meeting summaries, documents and software engineering | ANZ reported more than 7,000 emails drafted, 1,000 emails summarized and more than 1,000 meetings transcribed in one month, while describing material engineering gains without publishing a quantified productivity effect. | 65 / Moderate | Source |
| TBS-P187 | Procter & Gamble | Consumer goods | Algorithmic retailer-search advertising system – Automated search-ad bidding and optimization | P&G reported that its proprietary system automatically adjusted retailer search-ad buying every 15 minutes, increasing brand sales return by four times. | 78 / Moderate | Source |
| TBS-P188 | Procter & Gamble | Consumer goods / manufacturing | AI-enabled real-time vision inspection – Manufacturing-line product-quality inspection | P&G reported that real-time vision cameras and advanced algorithms analyzed more products for quality and drove productivity, but it did not publish a numerical defect, cost or throughput outcome. | 63 / Moderate | Source |
| TBS-P189 | Unilever Personal Care | Consumer goods / marketing | Personal Care AI Studio and Brand DNAi – Brand-asset creation, localisation and social-response production | Unilever reported faster asset creation, lower execution costs and greater responsiveness to social trends, with the AI studio live in four markets, but it published no numerical business outcome. | 73 / Moderate | Source |
| TBS-P190 | Unilever R&D | Consumer goods / research and development | AI-powered product simulations – Replacing physical trials during product research and innovation | Unilever reported that AI-powered simulations eliminated the need for multiple physical trials and accelerated innovation timescales, without publishing a numerical duration, cost or success-rate outcome. | 68 / Moderate | Source |
| TBS-P191 | L’Oréal | Beauty and consumer goods | L’Oréal GPT and generative-AI workforce tools – Companywide knowledge work and creativity | L’Oréal reported more than 65,000 employees trained in generative AI and more than 60,000 employees using L’Oréal GPT, but no quantified productivity, cost or revenue outcome. | 65 / Moderate | Source |
| TBS-P192 | UK Department for Business and Trade | Government / public administration | Microsoft 365 Copilot – Text-based knowledge work, meetings, drafting and information retrieval | DBT evaluated 1,000 Copilot licenses over three months and found high satisfaction and task-level time savings, particularly for text-heavy work and specialised user needs, but did not publish a single controlled aggregate productivity effect. | 68 / Moderate | Source |
| TBS-P193 | UK government evidence-review team | Government / policy research | Mixed generative-AI evidence-review tools – Rapid evidence review, study analysis and synthesis | The AI-assisted review was completed 23% faster overall and the literature-analysis phase took 56% less time, but the initial AI draft was less fluent and required more revisions than the human-only review. | 85 / Strong | Source |
| TBS-P194 | Representative UK workplace-task participants | Cross-industry work tasks | State-of-the-art large language model – Monitoring, technical drafting, work planning and information interpretation | Across four tasks, AI users scored 19% higher, completed work 25% faster and achieved 61% more points per minute; one planning task showed no statistically significant improvement. | 90 / Strong | Source |
| TBS-P195 | Large US materials R&D laboratory | Industrial research and development | AI materials-discovery technology – Candidate-material generation, testing and product innovation | AI-assisted scientists discovered 44% more materials, filed 39% more patents and produced 17% more downstream product innovations, while 82% reported lower job satisfaction. | 85 / Strong | Source |
| TBS-P196 | General Assembly | Education and professional training | MINT and IBM watsonx.ai – Marketing planning, budgeting, reporting and financial reconciliation | General Assembly reported 20% time savings in planning and reporting, 100% visibility into global budget pacing and one unified data model replacing fragmented regional taxonomies. | 80 / Strong | Source |
| TBS-P197 | IBM CIO organization | Enterprise IT | watsonx.data, watsonx.ai and governed data platform – Data-platform consolidation, governance and AI readiness | IBM reported US$5.3 million in savings after moving data from legacy systems and reducing redundancy, alongside 26% and 7.7% reductions in duplicated data in two domains. | 95 / Strong | Source |
| TBS-P198 | Claims Connection Group | Insurance services | IBM watsonx, workflow automation and RPA – Insurance claims interpretation and workflow orchestration | Claims Connection Group reported up to 70% faster claim resolution, a 30%–50% reduction in manual processing time and a 25%–35% reduction in operating costs. | 63 / Moderate | Source |
| TBS-P199 | Global pharmaceutical manufacturer | Pharmaceuticals | watsonx supply-chain risk orchestration – Assessing, triaging and escalating supply-chain risk events | The manufacturer reported more than 95% lower decision-cycle time, over 1,500 risk events assessed and no global event causing stock-out of top products in 2025. | 90 / Strong | Source |
| TBS-P200 | Digital Office Company | Document-management software | watsonx.ai and Watson Discovery – Document classification and metadata tagging | Digital Office Company reduced document-classification time from several minutes to 2.4 seconds during a six-week co-creation pilot. | 68 / Moderate | Source |
| TBS-P201 | Nubank | Financial technology / banking | GPT-4o customer assistant – Automating Tier 1 customer-service inquiries | Nubank’s assistant handled more than two million monthly chats, resolved about 55% of Tier 1 inquiries and reduced chat response time by 70%. | 85 / Strong | Source |
| TBS-P202 | Nubank | Financial technology / banking | Enterprise search and call-centre copilot – Internal policy search, chat summarization and next-reply suggestions | Nubank reported that its combined AI customer-service solutions helped resolve queries 2.3 times faster, while enterprise search reduced navigation across siloed information sources. | 63 / Moderate | Source |
| TBS-P203 | ABB Frosinone factory | Electrical manufacturing | AI-based visual defect-detection system – Automated circuit-breaker quality inspection | ABB reported that the AI defect-detection system reduced manual inspection time by 20%–40% and delivered a working proof of concept within one quarter. | 78 / Moderate | Source |
| TBS-P204 | Anonymous industrial steam and power plant | Industrial energy operations | ABB Ability OPTIMAX with AI forecasting – Energy-flow optimization and market-penalty reduction | ABB reported a 1.5% energy-cost reduction, 60% lower day-ahead penalties and 80% lower intraday penalties, with return on investment achieved within one year. | 85 / Strong | Source |
| TBS-P205 | Commonwealth Bank of Australia | Banking | CommBank Companion – Conversational personal and business financial guidance inside the banking app | CommBank began testing the assistant with employees and selected business customers, but published no numerical customer, financial, productivity or service outcome. | 40 / Limited | Source |
| TBS-P206 | DWF | Legal and professional services | Microsoft 365 Copilot – Legal research, presentations, client communications and service development | DWF reported completing one new-service development task in seven hours instead of seven days and another research-and-presentation workflow in under 90 minutes instead of two days. | 73 / Moderate | Source |
| TBS-P207 | Prague Airport | Aviation / airport operations | Microsoft 365 Copilot, Azure AI Services and Copilot Studio – Meetings, documentation, legal work and communications | Prague Airport reported that employees saved at least two hours per week and produced more consistent documentation after Copilot adoption. | 75 / Moderate | Source |
| TBS-P208 | Presidio | Digital services and technology consulting | Microsoft 365 Copilot and Azure OpenAI Service – Meetings, proposals, content, sales and project management | Presidio calculated about 1,200 employee hours saved monthly across 300 users, a 90% reduction in RFP work, six hours weekly saved by project managers and 70 new business opportunities linked to its Copilot programme. | 90 / Strong | Source |
| TBS-P209 | Commercial Bank of Dubai | Banking | Microsoft 365 Copilot – Meetings, documents, board administration and employee productivity | Commercial Bank of Dubai reported 39,000 hours saved annually, including 3,200 meeting summaries in one month and 15 hours saved monthly by the Board Secretary. | 85 / Strong | Source |
| TBS-P210 | Oriserve | Conversational AI software | Amazon Bedrock and multilingual generative-AI agents – Voice, chat and email customer engagement | Oriserve reported 50% lower cost to serve, 50% annual business growth, 90% speech-recognition accuracy and 60% fewer word errors than industry standards. | 75 / Moderate | Source |
| TBS-P211 | BPC | Payments software | Amazon Bedrock and Amazon Q Developer – Customer-service chatbot and software development | BPC reduced chatbot model costs by 66%; in a four-month pilot, eight developers generated more than 4,000 lines of code with a 46% acceptance rate and scanned 6,640 lines of Java. | 95 / Strong | Source |
| TBS-P212 | EXL | Data analytics and insurance services | Amazon Bedrock-powered LDS Underwriting Assist – Insurance underwriting assessment and document review | EXL reported reducing underwriting from several days to a few hours and estimated potential underwriting-cost savings of up to 80%; the product launched after 60 days of development. | 80 / Strong | Source |
| TBS-P213 | DoorDash | Delivery and logistics platform | Amazon Bedrock, Amazon Connect and Claude – Voice-operated self-service support for delivery workers | DoorDash built a live-test-ready generative-AI contact-centre solution in two months and later scaled it to hundreds of thousands of support calls daily, reporting material reductions in live-agent call volume. | 80 / Strong | Source |
| TBS-P214 | Viral Nation | Influencer marketing technology | Gemini and Vertex AI – Brand-safety analysis across influencer video and social content | Viral Nation reduced infrastructure costs by about 92% and in-house engineering resources by 70%, while monitoring 86,000 channels daily and scanning 15 years of content in 48 hours. | 85 / Strong | Source |
| TBS-P215 | Kapiche | Customer-experience analytics software | Gemini, Vertex AI and Google Kubernetes Engine – Conversation intelligence and AI infrastructure | Kapiche reported approximately 75% lower infrastructure-management overhead, 99.9%+ uptime and automatic handling of 10-fold traffic spikes across millions of customer interactions. | 85 / Strong | Source |
| TBS-P216 | Huge | Design and technology services | Gemini Enterprise and Gemini Enterprise Agent Platform – Client vetting, contracts, presentations and creative asset production | Huge reduced new-business intake research from hours or days to minutes and reported hours saved per request, while expanding automated video and creative services. | 68 / Moderate | Source |
| TBS-P217 | ALLO Fiber | Telecommunications | Gemini Enterprise, BigQuery and Looker – Conversational business analysis and reporting | ALLO Fiber reduced a complex analysis workflow from eight hours to 15 minutes, a reduction of more than 95%. | 80 / Strong | Source |
| TBS-P218 | Zoom | Collaboration software | Claude within Zoom AI Companion – Meeting summaries, meeting questions and collaborative documents | Zoom reported a 14% improvement in meeting-summary accuracy after adding Claude and integrated new Claude models into production within two weeks of release. | 85 / Strong | Source |
| TBS-P219 | Palo Alto Networks | Cybersecurity software | Claude on Google Cloud Vertex AI – Code generation, debugging, architecture explanation and developer onboarding | Palo Alto Networks reported a 20%–30% increase in feature-development velocity, onboarding time reduced from months to weeks and 70% faster integration tasks for junior developers in a pilot. | 90 / Strong | Source |
| TBS-P220 | Rox | Sales intelligence software | OpenAI-powered sales agents – Account research, sales preparation and engagement monitoring | Rox customers reported more than eight hours saved weekly per sales representative, 35% higher customer engagement and a two-fold increase in sales-accepted pipeline; Rox grew from zero to 25 enterprise accounts in seven months. | 90 / Strong | Source |
| TBS-P221 | Target | Retail | ChatGPT Enterprise and OpenAI APIs – Headquarters productivity, store processes, forecasting and digital experiences | Target reported that ChatGPT Enterprise had been rolled out to 18,000 headquarters employees, but the source did not publish a numerical productivity, cost, revenue or quality outcome. | 60 / Moderate | Source |
| TBS-P222 | OpenAI go-to-market organization | Artificial intelligence software | Internal GTM Assistant – Sales briefs, account research, Q&A, recaps and CRM support | OpenAI reported a 20% productivity lift for the average sales representative, equivalent to approximately one additional working day per week, with an average of 22 assistant messages weekly. | 85 / Strong | Source |
| TBS-P223 | Happiest Minds Technologies | Technology services | IBM watsonx.ai – Natural-language document search and business-chart creation | A two-month proof of concept produced a 30% productivity increase and reduced business-chart creation time by 15 hours per feature per month while searching 50 reports and documents monthly. | 85 / Strong | Source |
| TBS-P224 | SMS DataTech | Technology and cybersecurity services | IBM watsonx.ai – Cybersecurity intelligence collection and executive reporting | SMS DataTech reduced daily security-information collection from two hours to 10 minutes, a 92% reduction, and monthly executive-report creation from 40 hours to 10 hours. | 85 / Strong | Source |
| TBS-P225 | Savage Precision Fabrication | Aerospace and defence manufacturing | IBM watsonx Orchestrate – Quoting, production planning and procurement | Savage Precision reported a four-to-ten-fold increase in workforce productivity, quote creation reduced from hours to minutes and up to a 99% reduction in planning time. | 80 / Strong | Source |
| TBS-P226 | Booking.com | Travel technology | ServiceNow AI Agents – IT quality assurance, ticket classification and employee-service routing | Booking.com reported 97% correct ticket assignment, a 33% reduction in resolution time, quality-review coverage rising from 7% to 100% and quality scores increasing 30% over six months. | 95 / Strong | Source |
| TBS-P227 | Bell Canada | Telecommunications | ServiceNow AI Agents – Preparing, validating and routing high-stakes customer escalations | Bell Canada reported a 25% faster customer response time and 90% accuracy on AI-assisted escalation cases across thousands of cases annually. | 85 / Strong | Source |
| TBS-P228 | Fujitsu | Technology services | ServiceNow Otto – Employee search, knowledge creation and ticket handling | Fujitsu deployed Otto across 90,000 employees and projected a 50% improvement in search success, a 43% increase in knowledge creation and a 22% reduction in ticket-reassignment handling time. | 80 / Strong | Source |
| TBS-P229 | Engine | Travel technology | Agentforce – Autonomous handling of routine booking-cancellation requests | Engine reported a 15% reduction in average handle time and estimated that Agentforce would deliver US$2 million in annual cost savings. | 75 / Moderate | Source |
| TBS-P230 | Equitable | Financial services and insurance | Agentforce Marketing – Campaign production, segmentation and sweepstakes administration | Equitable reported a 75% reduction in campaign-build time, segmentation-list creation falling from 60 minutes to five minutes and a 75% reduction in sweepstakes administrative work. | 85 / Strong | Source |
| TBS-P231 | DLA Piper | Legal and professional services | Microsoft 365 Copilot – Legal research, content creation, data analysis and knowledge work | DLA Piper reported that early Copilot experiments saved participating employees up to 36 hours per week across content-creation and data-analysis workflows. | 73 / Moderate | Source |
| TBS-P232 | Kantar | Market research and consulting | Microsoft 365 Copilot – Research synthesis, meetings, documents and client-service knowledge work | Kantar reported 85% daily use among enabled employees and average savings of 1.8 hours per week, with some users reporting up to 10 hours weekly. | 80 / Strong | Source |
| TBS-P233 | AvePoint | Enterprise software | Microsoft 365 Copilot – Content production, software development and employee productivity | AvePoint reported reducing content-creation time by up to 66% and accelerating software-development workflows after companywide Copilot adoption. | 73 / Moderate | Source |
| TBS-P234 | Somerset Council | Local government | Microsoft 365 Copilot – Administrative work, documents, meetings and resident-service support | Somerset Council reported average savings of approximately 10 employee hours per month after adopting Copilot. | 75 / Moderate | Source |
| TBS-P235 | Thomson Reuters | Professional information and software | AWS Transform for .NET – Agentic modernisation of legacy enterprise applications | Thomson Reuters reported modernizing 1.5 million lines of code monthly, a four-fold increase in velocity, 30% lower costs, transformation reduced from months to one two-week sprint and 50% lower technical debt. | 95 / Strong | Source |
| TBS-P236 | Chime Financial | Financial technology | Amazon Bedrock call-summarization system – Customer-call notes, post-call summaries and service analytics | Chime reduced average call handling by 18 seconds, saved more than 250,000 hours annually, estimated approximately US$700,000 in annual efficiency gains and increased service NPS by five points. | 95 / Strong | Source |
| TBS-P237 | Sun Life | Insurance and asset management | Sun Life Asks on Amazon Bedrock – Internal enterprise search and question answering | Sun Life’s chatbot handled more than 10,000 queries weekly and over 600,000 queries in its first 11 months; employees reported saving several hours on selected research tasks. | 85 / Strong | Source |
| TBS-P238 | Vector Limited | Energy and infrastructure | AWS Transform for VMware – Agentic cloud migration, wave planning and infrastructure modernisation | Vector completed migration 34% faster, achieved 35% lower five-year total cost of ownership, automated about 60% of wave planning, improved team effectiveness by 30% and freed 17% more time for higher-value work. | 95 / Strong | Source |
| TBS-P239 | SIGNAL IDUNA | Insurance | Gemini and Vertex AI document-processing system – Insurance-document classification and case closure | A 20-person experiment found 30% lower processing time and case closure improving from 73% to almost 98%; the system processed approximately 2,000 documents across 600 tariff types with six-second average response time. | 95 / Strong | Source |
| TBS-P240 | Mattel | Consumer products and entertainment | Gemini and BigQuery customer-feedback analysis – Analysing product feedback and consumer insights | Mattel reported an estimated US$1 million in cost savings, 100-fold higher data-processing capacity and reduction of a feedback-analysis workflow from one month to one minute. | 85 / Strong | Source |
| TBS-P241 | Hiscox | Insurance | Vertex AI underwriting model – Automated quoting for complex insurance risks | Hiscox reduced complex-risk quoting from approximately three days to a few minutes using an AI-enhanced lead-underwriting model. | 75 / Moderate | Source |
| TBS-P242 | Wasco Berhad | Energy and industrial infrastructure | Gemini and Vertex AI safety-intelligence system – Global safety-observation analysis, risk categorisation and prediction | Wasco reduced safety-data analysis from weeks to minutes, saved thousands of employee hours and centralised insights from thousands of monthly observations across 13 countries. | 85 / Strong | Source |
| TBS-P243 | AdventHealth | Healthcare | OpenAI-powered administrative and clinical-support systems – Healthcare administration and employee workflows | AdventHealth reported an 80% reduction in time required for selected administrative tasks across a health system operating in nine states. | 73 / Moderate | Source |
| TBS-P244 | Dai Nippon Printing | Printing, information and manufacturing services | ChatGPT Enterprise – Companywide knowledge work, process automation and document handling | DNP reported 100% weekly active use among enabled users, measurable outcomes in 90% of use cases, up to 87% task automation and a ten-fold increase in processing volume for selected workflows. | 80 / Strong | Source |
| TBS-P245 | Oscar Health | Health insurance and healthcare technology | OpenAI-powered clinical documentation and claims-support tools – Clinical documentation and insurance claims escalation | Oscar reported 40% lower clinical-documentation time, 50% faster claims-escalation resolution and accuracy comparable with or better than humans; it expected automation of approximately 4,000 tickets monthly. | 90 / Strong | Source |
| TBS-P246 | IBM CIO and CISO organization | Enterprise technology | IBM Concert – Application-security risk prioritisation and vulnerability management | IBM reported analysing 874 applications in 24 hours, identifying 32% more vulnerabilities and reducing low-priority alerts by 67%. | 100 / Strong | Source |
| TBS-P247 | La Trobe University | Higher education and research | IBM watsonx.ai research assistant – Autism research synthesis and evidence analysis | La Trobe University reported analysing more than 100 autism research papers, building the AI component in under one week, reducing development from months to weeks and achieving six-fold time and 8.7-fold cost savings versus outsourcing. | 75 / Moderate | Source |
| TBS-P248 | ServiceNow | Enterprise software | Internal security and risk AI agents – Server patching, security incidents, risk assessments and phishing triage | ServiceNow reported 50% higher server-patch productivity, 85% faster security-incident closure, 67% less manual risk-assessment work and more than 1,700 analyst hours reclaimed monthly from phishing triage. | 90 / Strong | Source |
| TBS-P249 | Salesforce | Enterprise software | Agentforce on Salesforce.com – Website lead qualification, sales development and marketing pipeline | Salesforce reported 150 sales-development hours saved monthly, more than 30,000 new leads in 2025, 1.8-times higher lead conversion, approximately US$600,000 incremental marketing pipeline per 1,000 visitors and a US$20 million pipeline lift. | 95 / Strong | Source |
| TBS-P250 | Hero FinCorp | Financial services | Agentforce and Salesforce lending workflow – Loan application processing, dealer coordination and credit operations | Hero FinCorp reduced loan-approval turnaround from more than two days to under 30 minutes, reported an 80% turnaround reduction and expected 35% return over three years, 75% fewer handoffs and 37% fewer errors. | 90 / Strong | Source |
| TBS-P251 | Penda Health | Healthcare | AI Consult clinical copilot – Real-time clinical decision support during primary-care consultations | Across 39,849 visits, clinicians with AI Consult made 16% fewer diagnostic errors and 13% fewer treatment errors; independent physicians reviewed 5,666 randomly selected visits. | 95 / Strong | Source |
Appendix B. Scoring rubric and sampling controls
| Criterion | Points | Definition |
| Quantified result | 15 | Does the source publish a numerical outcome? |
| Baseline disclosure | 10 | Is there a usable comparator or pre-AI baseline? |
| Measurement period | 10 | Is the observation or evaluation window disclosed? |
| Deployment / sample size | 10 | Are users, tasks, records, calls, visits, documents, or similar denominators disclosed? |
| Methodology detail | 10 | Does the source explain how the result was calculated or evaluated? |
| Production maturity | 15 | Is the system in real production rather than only a demo or early pilot? |
| Business materiality | 10 | Is the outcome tied to a business, operational, clinical, risk, or financial metric? |
| Denominator / magnitude | 5 | Can the scale of the result be understood? |
| Adopter evidence | 10 | Is the adopter itself a direct evidence source or clearly identified? |
| Limitations disclosed | 5 | Does the source acknowledge caveats, uncertainty, mixed effects, or constraints? |
Sampling controls
- No single vendor ecosystem exceeds 10% of the final 250-case sample.
- Adopter-owned, financial, and independent evidence exceeds the minimum non-vendor share set during collection.
- Duplicate deployments were consolidated before freezing the dataset.
- Multiple cases from one organization were retained only when workflows were operationally distinct.
- Case-weighted and organization-balanced sensitivity results were compared to test concentration risk.
- Weak, adoption-only, failure, mixed, and null-result cases were retained rather than filtered out.
Appendix C. Validation and adjudication notes
A stratified 25-case second-pass AI validation recode was conducted after the dataset freeze. It is not an independent human review and does not satisfy the independent second-review requirement.
| Validation metric | Result |
| Cases recoded | 25 |
| Evidence-band agreement | 72% |
| Within ±5 points | 64% |
| Within ±10 points | 80% |
| Pearson score correlation | 0.813 |
| Quadratic weighted band kappa | 0.668 |
| Clear extraction corrections | 2 |
Two clear extraction issues were corrected in an adjudicated sensitivity dataset: LSEG’s source contained quantified cycle-time outcomes that the original row undercaptured, and Armanino’s source contained more explicit baseline and run-rate detail than originally coded.
Those two corrections changed the overall evidence-score average only slightly, supporting the conclusion that the dataset-level findings are not dependent on either case. Other larger scoring disagreements were not silently overwritten and remain pending independent adjudication.