05 / Quantitative ML · Full-stack
Stock Alpha
An S&P 500 momentum screen with a web app and an in-browser Monte Carlo portfolio simulator. The interesting part is what I removed: a 38-feature ML model, filing sentiment and fundamental screens were all tested blind and dropped when they failed. One feature survived.
24.2% vs 9.9%blind test, top 15 vs index
~500 stocksS&P 500 universe
SEC + Yahoo→ PostgreSQL
1,000 pathsMonte Carlo, 252 days
The signal
score = 6-month return (skip latest month) ÷ annualised volatility
The return runs from seven months ago to one month ago. The latest month is skipped because very short-term moves tend to reverse. Dividing by volatility (measured over about 126 trading days and annualised) rewards steady climbers over stocks that got there by lurching around. Profitability, growth and margins are shown next to each stock for context, but only the filter uses them.
What got cut, and why
| Tried | Result |
|---|---|
| ML model (38 features, XGBoost / RF) | Blind test AUC 0.50, no better than a coin flip |
| Fundamental screens ×3 | Each one lowered returns compared with momentum alone |
| Filing sentiment (VADER, TF-IDF + SVD) | Added noise, no signal |
| Macro-economic features | Made results worse |
| Other momentum variants | None beat plain 6-month return |
Bugs found by validating properly
- The SEC parser read only the first matching tag, so Apple, Nvidia and JPMorgan had no data.
- About 40% of stocks were silently dropped because of missing share counts.
- Filings were used about 45 days before they were public, which is look-ahead bias.