SEC fundamentals / annual forecasting

Forecasting Industry-Relative Revenue Growth

An expanding-window model that ranks U.S. firms against aligned industry peers.

2016–2024 holdouts · 25,221 company-years View source ↗

1

Forecasting question

Given information available by the end of fiscal year t−1, can a statistical learning model identify firms whose fiscal-year t revenue growth will exceed the aligned median growth of comparable SIC2 firms?

The estimand is cross-sectional and industry-relative. For firm i in prediction year t, the continuous target equals realized revenue growth less the leave-one-out median growth of aligned peers. A positive value defines the binary outcome.

2

Data and target construction

Annual financial statement variables are derived from SEC filings. Quarterly revenue facts are reconciled into fiscal-year trajectories and supplement the annual feature panel. Same-filing prior-year values are used where available to reduce inconsistencies introduced by later restatements. FRED series are retained for descriptive context but do not enter the final learner; SIC identifiers define peer sets but are not supplied as predictors.

Table 1. Study design and sample construction
PopulationU.S. public-company fiscal-year observations with eligible SEC financial statement data
Predictor timingFiscal year t−1 and earlier only
Peer definitionSame two-digit SIC industry; fiscal-period ends aligned within ±92 days
Peer statisticLeave-one-out median revenue growth; at least 20 aligned firms required
Gap control300–430 days between adjacent annual fiscal periods
Predictor set163 numeric firm variables after eligibility controls; no macro variables or industry indicators
EvaluationNine expanding-window annual holdouts, 2016–2024
Figure 1. Temporal separation of predictors and outcomes. The asterisk denotes a reconstructed fourth quarter where applicable. No fiscal-year t outcome information is available during model estimation.
System flowFrom filing to score
01 / Source

Filings establish the information set.

Annual 10-K statement facts supply accounting levels, ratios, and year-over-year changes. Reconciled quarterly revenue observations describe the within-year path.

3

Empirical design

Evaluation follows an expanding-window design. The 2016 model is estimated on 2011–2015 observations, and each subsequent evaluation adds the preceding year to the historical pool. Median imputation, first- and ninety-ninth-percentile winsorization, constant-column removal, and model fitting are repeated within each training window. The held-out year contributes no preprocessing statistics.

Expanding-window allocation by fiscal year

Training + internal validationActual test holdout

Purple cells are not final test observations. They form the expanding historical pool used for model fitting and the internal validation scores that govern early stopping. The red cell is the untouched annual test holdout from which the reported accuracy and ranking metrics are calculated.

Figure 2. Expanding-window evaluation design. For each prediction year, all eligible earlier fiscal years form the training and internal-validation pool; the diagonal red cell is then scored once as the actual out-of-sample test. Point to a purple cell to inspect the size of its historical window, or to a red cell to inspect holdout observations and realized accuracy, rank AUC, and top-decile precision. The grid also supports keyboard arrow navigation.

4

Out-of-sample results

Accuracy is the share of correct binary classifications at the 0.5 model-probability threshold. Precision is the share of predicted positive cases that realize a positive industry-relative target. AUC measures ranking discrimination across all positive-negative pairs. Top-decile precision is the realized positive share among the highest-ranked 10% of firms within each prediction year.

Pooled accuracy68.22%
Pooled positive precision69.20%
Mean annual rank AUC74.95%
Mean annual top-decile precision89.61%

Annual held-out performance

Fiscal years 2016–2024
Figure 3. Accuracy, rank-aggregation AUC, and top-decile precision by annual holdout. Values are computed only from observations withheld from estimation and preprocessing. Point to a marker to report the fiscal year, metric value, and number of holdout observations.

Precision by annual ranking threshold

Mean across nine holdout years
Figure 4. Mean annual precision as the retained ranking fraction narrows. Point to a bar to report the ranking threshold, mean annual precision, and mean number of company-year observations retained per annual holdout.
Table 2. Annual out-of-sample results
Prediction yearObservationsAccuracyRank AUCTop-decile precisionTop-decile observations
20162,67270.32%76.93%94.03%268
20172,57368.56%74.97%93.02%258
20182,78269.84%76.44%91.04%279
20192,76068.91%75.44%92.75%276
20202,79766.71%73.15%88.21%280
20212,89662.91%69.54%85.17%290
20223,07268.20%75.12%84.42%308
20232,90168.11%75.68%88.66%291
20242,76870.81%77.28%89.17%277
Mean / total25,22168.26%74.95%89.61%281 mean

The difference between pooled accuracy (68.22%) and mean annual accuracy (68.26%) reflects weighting: the pooled statistic weights company-year observations, whereas the annual mean gives each prediction year equal weight. Ranking performance is materially stronger than threshold classification performance, so the evidence is most informative for comparative annual ordering.

5

Variation across industries

Predictive separation is not uniform across peer groups. Among the 27 two-digit SIC industries with at least 200 held-out observations, business services and software records the highest pooled classification accuracy (74.22%), while holding and investment offices records the highest within-industry top-decile precision (97.44%). At the other end of the distribution, security and commodity brokers, wholesale trade, and depository institutions exhibit substantially weaker concentration precision.

SIC2 performance field27 peer groups
classification accuracy →
top-decile precision →
All eligible industries 27 SIC2 groups Marker size and depth reflect held-out sample size.
Distribution

Predictability varies across peer groups.

Each marker is a two-digit SIC industry. Horizontal position records pooled classification accuracy; vertical position records precision within the highest-scoring 10% of each industry-year cell.

Ranking separation

Holding and investment offices

The highest concentration precision is 97.44%: 190 of the 195 selected company-years realize positive industry-relative revenue growth.

Threshold classification

Business services and software

The largest eligible peer group also records the highest pooled classification accuracy, 74.22%, across 4,014 held-out company-years.

Lower separation

Security and commodity brokers

Within-industry top-decile precision falls to 66.67%. The dispersion indicates that the model's ranking evidence should not be assumed uniform across SIC2 groups.

Higher separation97.44%

Holding and investment offices
195 selected from 1,910 holdouts

Higher classification accuracy74.22%

Business services and software
4,014 holdouts across nine years

Lower concentration precision66.67%

Security and commodity brokers
69 selected from 646 holdouts

Industry-level predictive separation

27 SIC2 groups with at least 200 holdouts
Figure 5. Pooled classification accuracy and precision among the highest-scoring 10% within each industry-year cell. Marker area represents the number of held-out company-year observations. Dashed lines report the corresponding full-sample benchmarks. Point to an industry to inspect its SIC2 code, sample size, number selected, and both performance measures. Industry-level estimates are descriptive and differ in sampling uncertainty.

6

Model specification

Each component is a histogram gradient-boosting classifier with a learning rate of 0.05, a maximum of 300 iterations, 31 maximum leaf nodes, L2 regularization of 0.1, and early stopping. Within the historical pool defined above, fitting learns the recursive multivariate slices illustrated next. Five random seeds repeat this procedure on the same expanding training window.

Multivariate slicing

How recursive thresholds convert 163 firm signals into interaction profiles

The final learner does not begin with hand-written profile categories. Inside each expanding training window, training-only preprocessing retains 157–161 of 163 candidate firm predictors. A tree node selects one binned feature threshold at a time; nested nodes can then split the resulting branch using different predictors. The terminal leaves are therefore multivariate slices: firm profiles defined by combinations of annual levels, ratios, growth rates, quarterly dynamics, and missingness conditions.

Candidate inputs163 numeric firm features
Retained by window157–161 features
Tree complexity≤31 terminal leaves
Boosting path≤300 corrections · η 0.05
RegularizationL2 0.1 + early stopping
StabilityFive seeded fits
Multivariate slicingFrom thresholds to firm profiles
Firm profiles

Each observation enters as a profile, not a single ratio.

The learner receives up to 163 annual, quarterly, ratio, growth, and missingness features for the same firm-year. The two axes shown here are an explanatory projection of that higher-dimensional profile.

Histogram search

Continuous values are organized into candidate thresholds.

Histogram gradient boosting groups observed feature values into ordered bins. At each node, it searches the available features and thresholds for the split that most improves the loss within the current training window.

Multivariate slicing

One node is univariate; the resulting slice is multivariate.

A node may first divide firms by revenue momentum. A later node can divide only one branch by quarterly acceleration. Recursive one-feature decisions therefore create nonlinear interaction profiles without imposing a single linear response.

Boosted correction

Later trees concentrate on residual classification error.

Each tree contributes a regularized correction at a learning rate of 0.05. Early stopping is governed inside the historical training and validation pool; five seeded fits are then averaged before holdout classification and ranking.

Only after the five learners have been fitted are their outputs combined. Mean probabilities support threshold classification at 0.5; mean within-year percentile ranks support AUC and top-k precision. The rank score is a relative annual ordering, not a calibrated probability.

163 numeric firm predictorsCompany × prediction yearMacro variables and SIC identifiers excluded
HGB · seed 2026HGB · seed 2027HGB · seed 2028HGB · seed 2029HGB · seed 2030
Classification branchArithmetic mean probabilityThreshold at 0.5 → pooled accuracy and positive precision
Ranking branchMean within-year percentile rankRank AUC and top-k precision within each fiscal year
Figure 6. Final model architecture. Each seed-specific learner first constructs recursive multivariate slices during fitting. Threshold-based statistics then use the arithmetic mean of seed-specific probabilities, while ranking statistics use the arithmetic mean of seed-specific within-year percentile ranks.

7

Robustness analyses

Two descriptive analyses examine temporal weighting and annual distribution shift. These analyses characterize sensitivity of the fitted procedure; they do not identify causal mechanisms. Both three-dimensional figures support rotation, zooming, and point-level inspection.

Temporal-weight perturbations

Projected five-component simplex
Drag horizontally to rotate · hover to inspect
Figure 7. Dirichlet perturbation of five annual-history weights. Color records top-decile precision and marker size records AUC in the screening analysis. The reference vector [1,0,0,0,0] assigns all weight to the most recent fiscal year; its multi-seed evaluation yields 89.61% mean annual top-decile precision, compared with 88.55% for the highest-scoring smoothed specification. Drag horizontally to rotate, point to an observation to inspect its weights and metrics, or use Reset view to restore the initial camera.

Annual feature-distribution shift

First three principal components
Drag horizontally to rotate · hover to inspect
Figure 8. Principal-component representation of annual feature-distribution fingerprints based on means, medians, interquartile ranges, and missingness rates. The first three components account for 67.1% of observed annual distribution variation. Feature-shift distance is negatively associated with top-decile precision across the nine evaluated prediction years (Pearson r = −0.73; p = 0.024), an association that is descriptive rather than causal. Drag horizontally to rotate, point to a year to inspect its coordinates and subsequent evaluation metrics, or use Reset view to restore the initial camera.

8

Interpretation and limitations

The pooled classification results indicate moderate discrimination across the full company sample. The larger precision estimates in the upper ranking slices indicate that the specification is better suited to identifying a concentrated set of firms with comparatively high expected industry-relative revenue growth than to assigning equal evidential weight to every binary classification.

The rank-ensemble score expresses relative model ordering, not calibrated confidence. Where a binary classification is reported, the underlying probability may still indicate how far an observation lies from the 0.5 decision threshold, but this distance should not be interpreted as a validated probability of realized outperformance.

Limitations

  • SEC XBRL tag usage varies across issuers and over time, leaving residual accounting heterogeneity after canonicalization.
  • The target is relative to the observed peer sample and depends on SIC2 classification, fiscal-period alignment, and sample composition.
  • Delistings, late filings, mergers, restatements, and reporting coverage may affect the observed panel.
  • Nine annual holdouts provide limited evidence about performance under future macroeconomic or reporting regimes.
  • The study does not model valuation, risk, liquidity, transaction costs, or portfolio outcomes and is not an investment recommendation.

9

Reproducibility

The public repository contains the SEC, quarterly, target, and optional FRED data pipelines; the final five-seed model; the frozen feature manifest; annual evaluation outputs; robustness scripts; Wolfram Language source for the three-dimensional analyses; and automated tests.

Acknowledgements and references

Acknowledgements

This analysis depends on public financial-statement infrastructure maintained by the U.S. Securities and Exchange Commission and macroeconomic series distributed by the Federal Reserve Bank of St. Louis. The implementation also uses the open-source Python, scikit-learn, Plotly, and Three.js ecosystems. No external sponsor or institutional affiliation is asserted.

Data and software references

  1. U.S. Securities and Exchange Commission. Financial Statement Data Sets.
  2. U.S. Securities and Exchange Commission. EDGAR application programming interfaces.
  3. Federal Reserve Bank of St. Louis. FRED economic data.
  4. Occupational Safety and Health Administration. Standard Industrial Classification Manual.
  5. scikit-learn developers. Histogram Gradient Boosting Classifier.
  6. Plotly.js and Three.js. Interactive figure and WebGL rendering libraries.