SEC fundamentals / annual forecasting
Forecasting Industry-Relative Revenue Growth
An expanding-window model that ranks U.S. firms against aligned industry peers.
1
Forecasting question
Given information available by the end of fiscal year t−1, can a statistical learning model identify firms whose fiscal-year t revenue growth will exceed the aligned median growth of comparable SIC2 firms?
The estimand is cross-sectional and industry-relative. For firm i in prediction year t, the continuous target equals realized revenue growth less the leave-one-out median growth of aligned peers. A positive value defines the binary outcome.
2
Data and target construction
Annual financial statement variables are derived from SEC filings. Quarterly revenue facts are reconciled into fiscal-year trajectories and supplement the annual feature panel. Same-filing prior-year values are used where available to reduce inconsistencies introduced by later restatements. FRED series are retained for descriptive context but do not enter the final learner; SIC identifiers define peer sets but are not supplied as predictors.
| Population | U.S. public-company fiscal-year observations with eligible SEC financial statement data |
|---|---|
| Predictor timing | Fiscal year t−1 and earlier only |
| Peer definition | Same two-digit SIC industry; fiscal-period ends aligned within ±92 days |
| Peer statistic | Leave-one-out median revenue growth; at least 20 aligned firms required |
| Gap control | 300–430 days between adjacent annual fiscal periods |
| Predictor set | 163 numeric firm variables after eligibility controls; no macro variables or industry indicators |
| Evaluation | Nine expanding-window annual holdouts, 2016–2024 |
Available at estimation
Observed subsequently
Filings establish the information set.
Annual 10-K statement facts supply accounting levels, ratios, and year-over-year changes. Reconciled quarterly revenue observations describe the within-year path.
3
Empirical design
Evaluation follows an expanding-window design. The 2016 model is estimated on 2011–2015 observations, and each subsequent evaluation adds the preceding year to the historical pool. Median imputation, first- and ninety-ninth-percentile winsorization, constant-column removal, and model fitting are repeated within each training window. The held-out year contributes no preprocessing statistics.
Expanding-window allocation by fiscal year
Purple cells are not final test observations. They form the expanding historical pool used for model fitting and the internal validation scores that govern early stopping. The red cell is the untouched annual test holdout from which the reported accuracy and ranking metrics are calculated.
4
Out-of-sample results
Accuracy is the share of correct binary classifications at the 0.5 model-probability threshold. Precision is the share of predicted positive cases that realize a positive industry-relative target. AUC measures ranking discrimination across all positive-negative pairs. Top-decile precision is the realized positive share among the highest-ranked 10% of firms within each prediction year.
Annual held-out performance
Fiscal years 2016–2024Precision by annual ranking threshold
Mean across nine holdout years| Prediction year | Observations | Accuracy | Rank AUC | Top-decile precision | Top-decile observations |
|---|---|---|---|---|---|
| 2016 | 2,672 | 70.32% | 76.93% | 94.03% | 268 |
| 2017 | 2,573 | 68.56% | 74.97% | 93.02% | 258 |
| 2018 | 2,782 | 69.84% | 76.44% | 91.04% | 279 |
| 2019 | 2,760 | 68.91% | 75.44% | 92.75% | 276 |
| 2020 | 2,797 | 66.71% | 73.15% | 88.21% | 280 |
| 2021 | 2,896 | 62.91% | 69.54% | 85.17% | 290 |
| 2022 | 3,072 | 68.20% | 75.12% | 84.42% | 308 |
| 2023 | 2,901 | 68.11% | 75.68% | 88.66% | 291 |
| 2024 | 2,768 | 70.81% | 77.28% | 89.17% | 277 |
| Mean / total | 25,221 | 68.26% | 74.95% | 89.61% | 281 mean |
The difference between pooled accuracy (68.22%) and mean annual accuracy (68.26%) reflects weighting: the pooled statistic weights company-year observations, whereas the annual mean gives each prediction year equal weight. Ranking performance is materially stronger than threshold classification performance, so the evidence is most informative for comparative annual ordering.
5
Variation across industries
Predictive separation is not uniform across peer groups. Among the 27 two-digit SIC industries with at least 200 held-out observations, business services and software records the highest pooled classification accuracy (74.22%), while holding and investment offices records the highest within-industry top-decile precision (97.44%). At the other end of the distribution, security and commodity brokers, wholesale trade, and depository institutions exhibit substantially weaker concentration precision.
Predictability varies across peer groups.
Each marker is a two-digit SIC industry. Horizontal position records pooled classification accuracy; vertical position records precision within the highest-scoring 10% of each industry-year cell.
Holding and investment offices
The highest concentration precision is 97.44%: 190 of the 195 selected company-years realize positive industry-relative revenue growth.
Business services and software
The largest eligible peer group also records the highest pooled classification accuracy, 74.22%, across 4,014 held-out company-years.
Security and commodity brokers
Within-industry top-decile precision falls to 66.67%. The dispersion indicates that the model's ranking evidence should not be assumed uniform across SIC2 groups.
Holding and investment offices
195 selected from 1,910 holdouts
Business services and software
4,014 holdouts across nine years
Security and commodity brokers
69 selected from 646 holdouts
Industry-level predictive separation
27 SIC2 groups with at least 200 holdouts6
Model specification
Each component is a histogram gradient-boosting classifier with a learning rate of 0.05, a maximum of 300 iterations, 31 maximum leaf nodes, L2 regularization of 0.1, and early stopping. Within the historical pool defined above, fitting learns the recursive multivariate slices illustrated next. Five random seeds repeat this procedure on the same expanding training window.
How recursive thresholds convert 163 firm signals into interaction profiles
The final learner does not begin with hand-written profile categories. Inside each expanding training window, training-only preprocessing retains 157–161 of 163 candidate firm predictors. A tree node selects one binned feature threshold at a time; nested nodes can then split the resulting branch using different predictors. The terminal leaves are therefore multivariate slices: firm profiles defined by combinations of annual levels, ratios, growth rates, quarterly dynamics, and missingness conditions.
Each observation enters as a profile, not a single ratio.
The learner receives up to 163 annual, quarterly, ratio, growth, and missingness features for the same firm-year. The two axes shown here are an explanatory projection of that higher-dimensional profile.
Continuous values are organized into candidate thresholds.
Histogram gradient boosting groups observed feature values into ordered bins. At each node, it searches the available features and thresholds for the split that most improves the loss within the current training window.
One node is univariate; the resulting slice is multivariate.
A node may first divide firms by revenue momentum. A later node can divide only one branch by quarterly acceleration. Recursive one-feature decisions therefore create nonlinear interaction profiles without imposing a single linear response.
Later trees concentrate on residual classification error.
Each tree contributes a regularized correction at a learning rate of 0.05. Early stopping is governed inside the historical training and validation pool; five seeded fits are then averaged before holdout classification and ranking.
Only after the five learners have been fitted are their outputs combined. Mean probabilities support threshold classification at 0.5; mean within-year percentile ranks support AUC and top-k precision. The rank score is a relative annual ordering, not a calibrated probability.
7
Robustness analyses
Two descriptive analyses examine temporal weighting and annual distribution shift. These analyses characterize sensitivity of the fitted procedure; they do not identify causal mechanisms. Both three-dimensional figures support rotation, zooming, and point-level inspection.
Temporal-weight perturbations
Projected five-component simplex[1,0,0,0,0] assigns all weight to the most recent fiscal year; its multi-seed evaluation yields 89.61% mean annual top-decile precision, compared with 88.55% for the highest-scoring smoothed specification. Drag horizontally to rotate, point to an observation to inspect its weights and metrics, or use Reset view to restore the initial camera.Annual feature-distribution shift
First three principal components8
Interpretation and limitations
The pooled classification results indicate moderate discrimination across the full company sample. The larger precision estimates in the upper ranking slices indicate that the specification is better suited to identifying a concentrated set of firms with comparatively high expected industry-relative revenue growth than to assigning equal evidential weight to every binary classification.
The rank-ensemble score expresses relative model ordering, not calibrated confidence. Where a binary classification is reported, the underlying probability may still indicate how far an observation lies from the 0.5 decision threshold, but this distance should not be interpreted as a validated probability of realized outperformance.
Limitations
- SEC XBRL tag usage varies across issuers and over time, leaving residual accounting heterogeneity after canonicalization.
- The target is relative to the observed peer sample and depends on SIC2 classification, fiscal-period alignment, and sample composition.
- Delistings, late filings, mergers, restatements, and reporting coverage may affect the observed panel.
- Nine annual holdouts provide limited evidence about performance under future macroeconomic or reporting regimes.
- The study does not model valuation, risk, liquidity, transaction costs, or portfolio outcomes and is not an investment recommendation.
9
Reproducibility
The public repository contains the SEC, quarterly, target, and optional FRED data pipelines; the final five-seed model; the frozen feature manifest; annual evaluation outputs; robustness scripts; Wolfram Language source for the three-dimensional analyses; and automated tests.
Acknowledgements and references
Acknowledgements
This analysis depends on public financial-statement infrastructure maintained by the U.S. Securities and Exchange Commission and macroeconomic series distributed by the Federal Reserve Bank of St. Louis. The implementation also uses the open-source Python, scikit-learn, Plotly, and Three.js ecosystems. No external sponsor or institutional affiliation is asserted.
Data and software references
- U.S. Securities and Exchange Commission. Financial Statement Data Sets.
- U.S. Securities and Exchange Commission. EDGAR application programming interfaces.
- Federal Reserve Bank of St. Louis. FRED economic data.
- Occupational Safety and Health Administration. Standard Industrial Classification Manual.
- scikit-learn developers. Histogram Gradient Boosting Classifier.
- Plotly.js and Three.js. Interactive figure and WebGL rendering libraries.