Methodology Cards & Mathematical Specifications
In accordance with the project’s statistical transparency rules, every calculation and decomposition model is documented below with its exact mathematical formula, parameter definitions, core accounting assumptions, limitations, and validation status.
Proportional Disaggregation of Broad SUT Totals
\text{AI}_{g, t} = \text{BroadValue}_{g, t} \times s_{g, t}, \quad \text{NonAI}_{g, t} = \text{BroadValue}_{g, t} \times (1 - s_{g, t})Applies an evidence-based or scenario-tested AI share ($s_g \in [0, 1]$) to a published broad CPA product total to separate AI output from non-AI baseline output.
Parameter Definitions
| Parameter | Symbol | Description | Documented Default |
|---|---|---|---|
| Broad Value | BroadValue | Published ONS Supply & Use total for product group g in year t (£m) | Published SUT |
| Base AI Share | s_base | Central baseline AI share assumption based on survey / industry benchmarking | 16.5% (CPA J62) |
| Low Bound Share | s_low | Conservative lower bound AI share parameter | 9.0% (CPA J62) |
| High Bound Share | s_high | Upper sensitivity bound AI share parameter | 28.0% (CPA J62) |
- •The AI share is applied uniformly across the product transaction categories unless separate matrices exist.
- •Mathematical identity: AI Component + Non-AI Component strictly equals Published Broad Total.
- •Sensitivity to choice of share parameter requires interval testing rather than point precision.
- •A business count proportion is not automatically equivalent to a turnover/output proportion.
Firm-Level Modelled Disaggregation with Revenue Attribution
\text{Expected AI}_{i} = \text{Turnover}_{i} \times p_{\text{AI}, i} \times r_{\text{AI}, i}, \quad \text{AI}_{g} = \sum_{i \in g} \text{Expected AI}_{i}Decomposes enterprise revenue by explicitly separating the probability that a firm is AI-active ($p_{AI}$) from the proportion of its revenue attributable to AI ($r_{AI}$).
Parameter Definitions
| Parameter | Symbol | Description | Documented Default |
|---|---|---|---|
| Firm AI Probability | p_AI | Estimated probability of AI activity from text classification / filing features | 0.45 |
| Revenue Attribution Ratio | r_AI | Estimated share of firm revenue derived from AI products/services | 0.35 |
- •Firm AI probability and revenue attribution are distinct parameters; being 100% likely an AI company does not mean 100% of revenue is AI.
- •Diversified tech conglomerates have lower $r_{AI}$ than dedicated boutique AI labs.
- •Requires financial microdata or representative domain sampling to calibrate $r_{AI}$.
- •Small firms may have higher volatility in revenue attribution over time.
Hierarchical Tiered Hybrid Disaggregation
\text{AI}_{g} = V_{\text{direct}} + (\text{Residual}_1 \times s_{\text{prop}}) + (\text{Residual}_2 \times w_{\text{model}})Prioritizes direct observed monetary accounting values (Tier 1), allocates verified proportional shares for well-defined segments (Tier 2), and applies modelled estimation only to residual gaps (Tier 3).
Parameter Definitions
| Parameter | Symbol | Description | Documented Default |
|---|---|---|---|
| Direct Observed Tier | V_direct | Audited AI revenue/output from dedicated AI producers | £4.5bn |
| Proportional Tier Share | s_prop | Survey-calibrated share on remaining residual volume | 12.0% |
| Modelled Residual Weight | w_model | Statistical gap-filling weight for unmeasured segments | 6.0% |
- •Direct observed values are subtracted from the broad denominator before secondary shares are applied to avoid double counting.
- •Higher tiers carry higher evidentiary quality weights.
- •Requires maintaining tier boundaries and explicit evidence hierarchy logs.
Multi-Label Business Text Classification & TF-IDF Logistic Engine
P(\text{AI} \mid x) = \sigma\left(\sum_{j=1}^{K} w_j \cdot x_j + b\right) = \frac{1}{1 + e^{-(\mathbf{w}^T \mathbf{x} + b)}}Supervised classification of business activity text into 13 ONS Table 3 multi-label categories and binary AI relevance, with calibrated explainability feature weights.
Parameter Definitions
| Parameter | Symbol | Description | Documented Default |
|---|---|---|---|
| Vocabulary Weights | w_j | Calibrated logistic regression coefficient for term j | Trained on UK Corpus |
| Classification Threshold | Threshold | Decision boundary for binary AI relevance flag | 0.50 |
- •Business descriptions contain domain-specific vocabulary indicative of AI engineering versus non-technical adoption.
- •Human review status overrides algorithmic prediction for official register curation.
- •Susceptible to vocabulary drift as AI marketing buzzwords proliferate in non-technical sectors.
System of National Accounts (SNA 2008 / ESA 2010) Asset Boundary Evaluation
\text{GFCF}_{\text{own-account}} = \text{Labour Costs} + \text{Intermediate Inputs} + \text{Capital Services}Deterministic rule evaluation determining whether AI expenditure represents Gross Fixed Capital Formation (intangible IP asset AN.1173/AN.1171) or Intermediate Consumption (P.2).
Parameter Definitions
| Parameter | Symbol | Description | Documented Default |
|---|---|---|---|
| Service Life | T | Expected economic utility in production (> 1 year required for GFCF) | > 1 year |
| Economic Ownership | Owner | Party entitled to benefits and accepting operating risks | UK Resident |
- •Own-account software developed for internal use is valued at sum of production costs (SNA §10.137).
- •Cloud compute API fees without intellectual property asset transfer represent intermediate consumption.
- •Requires inspection of corporate contracts and economic ownership terms.