← All projects
05 / Quantitative ML · Full-stack

Stock Alpha

An S&P 500 momentum screen with a web app and an in-browser Monte Carlo portfolio simulator. The interesting part is what I removed: a 38-feature ML model, filing sentiment and fundamental screens were all tested blind and dropped when they failed. One feature survived.

24.2% vs 9.9%blind test, top 15 vs index
~500 stocksS&P 500 universe
SEC + Yahoo→ PostgreSQL
1,000 pathsMonte Carlo, 252 days
SEC EDGAR7 numbers / quarter Yahoo Finance2 yrs daily closes PostgreSQLstock_alpha.py FilterROA > 0mcap ≥ $300M Scoreret(t−7m → t−1m)÷ annual vol Ranktop 15 = BUY Flask API + web appGunicorn · Nginx Monte CarloGBM, in the browser Refreshed quarterly by cron: 0 2 1 1,4,7,10 *

The signal

score = 6-month return (skip latest month) ÷ annualised volatility

The return runs from seven months ago to one month ago. The latest month is skipped because very short-term moves tend to reverse. Dividing by volatility (measured over about 126 trading days and annualised) rewards steady climbers over stocks that got there by lurching around. Profitability, growth and margins are shown next to each stock for context, but only the filter uses them.

What got cut, and why

TriedResult
ML model (38 features, XGBoost / RF)Blind test AUC 0.50, no better than a coin flip
Fundamental screens ×3Each one lowered returns compared with momentum alone
Filing sentiment (VADER, TF-IDF + SVD)Added noise, no signal
Macro-economic featuresMade results worse
Other momentum variantsNone beat plain 6-month return

Bugs found by validating properly

  • The SEC parser read only the first matching tag, so Apple, Nvidia and JPMorgan had no data.
  • About 40% of stocks were silently dropped because of missing share counts.
  • Filings were used about 45 days before they were public, which is look-ahead bias.

Blind test: 5 snapshots, 2021–2025

Top-15 momentum
24.2%
Whole index
9.9%

Average 6-month return, top 15 held for six months. Beat the index in 3 of 5 snapshots. 2025 (the AI rally) carries most of the average, and without it the edge is about +5%. No transaction costs. These numbers are for plain momentum. The current version adds the profitability filter and volatility adjustment, which haven't been blind-tested yet.

Try it: the portfolio Monte Carlo

1,000 simulations
5th pct (bad)–
Median–
95th pct (good)–

Same method as the site: geometric Brownian motion, 252 trading days, 1,000 runs, starting from €10,000. Not financial advice.