Few industries generate data like telecoms. Every call, session and handover leaves a record; every cell, router and core function emits counters, alarms and logs; every customer has a contract, a device, a usage pattern, a bill and a contact history. In principle this makes operators natural AI leaders. In practice, the data is fragmented across dozens of BSS and OSS stacks, defined differently in each, and often constrained by where it may be stored and who may use it.
The consequence shows up in the results. In McKinsey's research, 30% of telecom executives cite limitations in their data as a core inhibitor of AI impact at scale, and 45% say data is the core inhibitor they foresee for scaling AI agents1. The same work describes a developer-copilot pilot whose 25% to 40% productivity gain fell to less than 5% at scale, with incomplete or inconsistent data foundations among the reasons1. A year later, only 57% of telcos reported scaling GenAI across multiple domains, virtually unchanged, and only 12% had captured sizable impact2.
The telecom data estate
This is not a new problem. When TM Forum asked operators in 2020 to rate how well they used their data, the average score from 106 responses across 46 operators was 53 out of 10012. Four families of data matter most for AI, and each has its own quality problem.
- Usage records (CDRs and xDRs). The financial truth of the network, feeding billing, revenue assurance, fraud and churn models. Quality issues are duplicates, gaps between mediation and billing, and inconsistent product codes.
- Network telemetry and probe data. Counters, alarms, traces and flow data from RAN, transport and core. Ericsson reports global mobile network traffic exceeded 220 EB a month in Q2 2026, up 23% year on year4; 4G RANs exposed hundreds of KPIs where 5G exposes thousands5. Volume, timeliness and vendor-specific semantics are the challenges.
- Customer and commercial data. Accounts, contracts, product holdings, interactions and payments, often duplicated across fixed, mobile and enterprise stacks with no single customer identifier.
- Unstructured data. Call transcripts, chat logs, engineer notes and contracts, the raw material for GenAI and the least governed of all.
The scale of the telecom data estate
Selected indicators of data volume and complexity
| Indicator | Figure | Source |
|---|---|---|
| Global mobile network data traffic, Q2 2026 | More than 220 EB per month, up 23% year on year | Ericsson |
| Data generated daily by one large European operator group | Around 1 PB | TM Forum Inform |
| RAN key performance indicators, 4G vs 5G | Hundreds vs thousands | TM Forum Inform |
| Operators' self-rated effectiveness in using their data (2020) | 53 out of 100 (106 responses, 46 operators) | TM Forum |
Note: Sources: [3], [4], [5], [12].
Source: Ericsson, “Mobile network traffic Q2 2026 (Mobility Report data and forecasts)” (2026)
One large European operator group reports generating around a petabyte of data every day, and has created an internal marketplace for data products; it also notes that some data, such as network topology and CDRs, cannot be exported outside the country3. That combination (enormous volume, reusable products, hard sovereignty constraints) is typical.
From lakes to data products
McKinsey observes that while some telcos moved early to build data products and digital twins of domains such as network and call centre, most have only started to use the impetus of GenAI to move to hybrid lakehouse architectures with structured data products that are curated and reusable across use cases1. The distinction matters. A lake stores data; a data product has an owner, a contract, quality tests, documentation and a known consumer. Cross-industry evidence shows the cost of skipping this step: Gartner found 63% of organisations did not have, or were unsure whether they had, the right data management practices for AI, and predicts organisations will abandon 60% of AI projects unsupported by AI-ready data through 20266.
Meaning is the missing layer
AI agents need more than clean tables; they need to know what a 'customer', an 'active subscriber' or a 'dropped call' means, and which system is authoritative. The industry has long had a shared vocabulary in TM Forum's Information Framework (SID), which provides a reference data model and common vocabulary across market and sales, customer, product, service, resource, partner and enterprise domains9. Leading operators are now investing in knowledge frameworks, or semantic data layers, as a foundational layer of AI-native architecture, though McKinsey cautions that ontology-driven approaches can be costly and complex to design, govern and update at enterprise scale2.
The pragmatic answer is to build the semantic layer incrementally, anchored to a standard model, covering the entities that the first priority use cases need, and to publish it as a governed product that both people and agents query. Lineage then records how each metric was derived, so an agent's recommendation can be traced back to source records, which is increasingly a governance requirement rather than a nice-to-have.
Privacy and consent by design
Telecom data is among the most sensitive consumers generate: it reveals location, relationships and behaviour. Regulators treat it accordingly. In 2024 the FCC fined four US wireless carriers nearly USD 200m for sharing access to customers' location data without the affirmative, express consent that section 222 of the Communications Act requires10. Operators feel the constraint internally too: in a 2025 cross-industry study, 65% of executives said data privacy rules limit their ability to use AI for personalisation, while 54% of consumers reported declining trust in companies' use of their data11. In TM Forum research, 80% of operators cited privacy and security as a top GenAI risk8.
Operators see data risk as the leading GenAI risk
Share of operators citing each concern, TM Forum GenAI survey, % (%)
Note: TM Forum research published December 2023.
The implication is that consent, purpose and residency must be attributes of the data itself, carried through the pipeline and enforced at query time, so that a churn model or care agent can only use what the customer has agreed to, in the jurisdiction where it is allowed. Retrofitting this after models are live is far more expensive than designing it in.
Gartner's prediction that at least 30% of GenAI projects would be abandoned after proof of concept by the end of 2025, with poor data quality among the causes, is a useful warning7. For operators, the order of work is clear: data products first, meaning second, agents third.