ReportTransformation

Data before agents: the foundations telecom AI actually runs on

Operators sit on some of the richest data in any industry, yet data is the barrier executives most often cite to scaling AI agents. The fix is less about new platforms and more about governed data products, a shared semantic layer, lineage and consent built in from the start.

9 min read By · Report
45%
of telecom executives surveyed say data is the core inhibitor they foresee for scaling AI agents1

Key takeaways

  • Data is the binding constraint: 45% of telecom executives name it as the core inhibitor to scaling AI agents, and 30% already cite data limitations as a barrier to impact at scale1.
  • Scale is stalling: only 57% of telcos report scaling GenAI use cases across multiple domains, and only 12% have captured sizable impact2.
  • The data estate is vast and growing, with global mobile traffic above 220 EB a month and 5G networks exposing thousands of KPIs rather than hundreds45.
  • Leaders are building curated data products and semantic layers that give AI agents consistent meaning, lineage and consent, rather than pointing models at raw lakes12.

Few industries generate data like telecoms. Every call, session and handover leaves a record; every cell, router and core function emits counters, alarms and logs; every customer has a contract, a device, a usage pattern, a bill and a contact history. In principle this makes operators natural AI leaders. In practice, the data is fragmented across dozens of BSS and OSS stacks, defined differently in each, and often constrained by where it may be stored and who may use it.

The consequence shows up in the results. In McKinsey's research, 30% of telecom executives cite limitations in their data as a core inhibitor of AI impact at scale, and 45% say data is the core inhibitor they foresee for scaling AI agents1. The same work describes a developer-copilot pilot whose 25% to 40% productivity gain fell to less than 5% at scale, with incomplete or inconsistent data foundations among the reasons1. A year later, only 57% of telcos reported scaling GenAI across multiple domains, virtually unchanged, and only 12% had captured sizable impact2.

The telecom data estate

This is not a new problem. When TM Forum asked operators in 2020 to rate how well they used their data, the average score from 106 responses across 46 operators was 53 out of 10012. Four families of data matter most for AI, and each has its own quality problem.

  • Usage records (CDRs and xDRs). The financial truth of the network, feeding billing, revenue assurance, fraud and churn models. Quality issues are duplicates, gaps between mediation and billing, and inconsistent product codes.
  • Network telemetry and probe data. Counters, alarms, traces and flow data from RAN, transport and core. Ericsson reports global mobile network traffic exceeded 220 EB a month in Q2 2026, up 23% year on year4; 4G RANs exposed hundreds of KPIs where 5G exposes thousands5. Volume, timeliness and vendor-specific semantics are the challenges.
  • Customer and commercial data. Accounts, contracts, product holdings, interactions and payments, often duplicated across fixed, mobile and enterprise stacks with no single customer identifier.
  • Unstructured data. Call transcripts, chat logs, engineer notes and contracts, the raw material for GenAI and the least governed of all.
Exhibit 1

The scale of the telecom data estate

Selected indicators of data volume and complexity

IndicatorFigureSource
Global mobile network data traffic, Q2 2026More than 220 EB per month, up 23% year on yearEricsson
Data generated daily by one large European operator groupAround 1 PBTM Forum Inform
RAN key performance indicators, 4G vs 5GHundreds vs thousandsTM Forum Inform
Operators' self-rated effectiveness in using their data (2020)53 out of 100 (106 responses, 46 operators)TM Forum

Note: Sources: [3], [4], [5], [12].

Source: Ericsson, “Mobile network traffic Q2 2026 (Mobility Report data and forecasts)” (2026)

One large European operator group reports generating around a petabyte of data every day, and has created an internal marketplace for data products; it also notes that some data, such as network topology and CDRs, cannot be exported outside the country3. That combination (enormous volume, reusable products, hard sovereignty constraints) is typical.

From lakes to data products

McKinsey observes that while some telcos moved early to build data products and digital twins of domains such as network and call centre, most have only started to use the impetus of GenAI to move to hybrid lakehouse architectures with structured data products that are curated and reusable across use cases1. The distinction matters. A lake stores data; a data product has an owner, a contract, quality tests, documentation and a known consumer. Cross-industry evidence shows the cost of skipping this step: Gartner found 63% of organisations did not have, or were unsure whether they had, the right data management practices for AI, and predicts organisations will abandon 60% of AI projects unsupported by AI-ready data through 20266.

Meaning is the missing layer

AI agents need more than clean tables; they need to know what a 'customer', an 'active subscriber' or a 'dropped call' means, and which system is authoritative. The industry has long had a shared vocabulary in TM Forum's Information Framework (SID), which provides a reference data model and common vocabulary across market and sales, customer, product, service, resource, partner and enterprise domains9. Leading operators are now investing in knowledge frameworks, or semantic data layers, as a foundational layer of AI-native architecture, though McKinsey cautions that ontology-driven approaches can be costly and complex to design, govern and update at enterprise scale2.

The pragmatic answer is to build the semantic layer incrementally, anchored to a standard model, covering the entities that the first priority use cases need, and to publish it as a governed product that both people and agents query. Lineage then records how each metric was derived, so an agent's recommendation can be traced back to source records, which is increasingly a governance requirement rather than a nice-to-have.

Privacy and consent by design

Telecom data is among the most sensitive consumers generate: it reveals location, relationships and behaviour. Regulators treat it accordingly. In 2024 the FCC fined four US wireless carriers nearly USD 200m for sharing access to customers' location data without the affirmative, express consent that section 222 of the Communications Act requires10. Operators feel the constraint internally too: in a 2025 cross-industry study, 65% of executives said data privacy rules limit their ability to use AI for personalisation, while 54% of consumers reported declining trust in companies' use of their data11. In TM Forum research, 80% of operators cited privacy and security as a top GenAI risk8.

Exhibit 2

Operators see data risk as the leading GenAI risk

Share of operators citing each concern, TM Forum GenAI survey, % (%)

Note: TM Forum research published December 2023.

Source: TM Forum, “Telcos unanimous on power of GenAI but implementation questions remain, TM Forum research finds” (2023)

The implication is that consent, purpose and residency must be attributes of the data itself, carried through the pipeline and enforced at query time, so that a churn model or care agent can only use what the customer has agreed to, in the jurisdiction where it is allowed. Retrofitting this after models are live is far more expensive than designing it in.

Gartner's prediction that at least 30% of GenAI projects would be abandoned after proof of concept by the end of 2025, with poor data quality among the causes, is a useful warning7. For operators, the order of work is clear: data products first, meaning second, agents third.

For executives

What this means for your operator

  1. Define five to seven priority data products (customer, product holding, usage, network performance, interactions) with named owners, contracts and quality SLAs.
  2. Build a semantic layer anchored to TM Forum SID for the entities your first AI use cases need, and make it queryable by agents as well as analysts.
  3. Instrument lineage end to end so every AI output can be traced to source records and transformations.
  4. Attach consent, purpose and residency to data at ingestion and enforce them at query time, including for unstructured transcripts.
  5. Stop funding AI use cases that build bespoke pipelines; require reuse of governed data products as a funding condition.
Put it to work

How DaasLabs can help

DaasLabs Data Fabric Framework reference architecture

Explore the framework

Data normalisation across BSS, OSS and network sources

See UNIFY data normalisation

Governed, auditable data with lineage

See compliance & data lineage

Data and AI maturity assessment

Take the maturity assessment

Sources

  1. 1
    Scaling the AI-native telco (opens in a new tab) McKinsey & Company, 27 February 2025
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7
  8. 8
  9. 9
  10. 10
  11. 11
    AI and the empathy gap: 2025 CX annual insights report (opens in a new tab) US operator business unit (cross-industry survey), 13 August 2025
  12. 12

Figures are drawn from the cited public sources. Opinions labelled “DaasLabs point of view” are our own.

Where this fits in the story

From connectivity provider to intelligent, AI-native operator

This piece is chapter 3 of 6: the foundation. A governed telecom data platform and a 12-layer target architecture, built once and reused.

  1. 01 The pressure 2 insights
  2. 02 The value chain 3 insights
  3. 03 The foundation 2 insights
  4. 04 The proof 2 insights
  5. 05 The workforce 3 insights
  6. 06 The journey Tools
Stay informed

Get new telecom insights in your inbox

New perspectives on AI, data and transformation in telecom — a few times a month. Browse all insights.

AI
AI Analyst

I'm the DaasLabs AI Analyst for the telecom demo platform. I can help with:

  • Revenue assurance & CDR reconciliation
  • Fraud: SIM swap, SIM box, IRSF and Wangiri
  • Churn, customer and network analytics
  • Executive briefings across the accelerators

Answers are generated from the demo data.