Fordham · CISC 5352, ML in Finance · with Patricia Angeles · Spring 2026
FinBERT Sentiment vs. Price-Based ML
An NLP pipeline over 480 SEC filings, tested against price-based classifiers to see whether filing sentiment predicts post-earnings returns. It barely does, which is consistent with semi-strong market efficiency.

- text chunks classified from 480 SEC 10-K/10-Q filings
- ~27,000text chunks classified from 480 SEC 10-K/10-Q filings
- sentiment–price agreement, below the 50% baseline
- 44.9%sentiment–price agreement, below the 50% baseline
- correlation between sentiment and abnormal return
- r = −0.043correlation between sentiment and abnormal return
The question
When a company files its 10-K or 10-Q, does the tone of the filing tell you anything about how the stock will move after earnings? And does it agree with what a model trained purely on price data predicts?
This was a two-person project with Patricia Angeles.
Part 1: Sentiment from SEC filings
We built an NLP pipeline over 480 SEC 10-K and 10-Q filings from 24 S&P 500 companies across eight sectors (2021–2025):
- Extraction. Pulled filings from EDGAR and kept only the sentiment-bearing sections (business overview, risk factors, MD&A), dropping financial tables and exhibits.
- Chunking. Cleaned the text and split it into sentence-aware, overlapping 128-token windows sized for FinBERT. Chunks that were mostly numbers were dropped. Each section was chunked on its own so risk-factor language never bled into MD&A.
- Classification. Ran zero-shot FinBERT on about 27,000 chunks, then aggregated them into 2,400 filing-level sentiment labels using confidence-weighted majority voting.
Part 2: Price-based prediction
In parallel, we engineered 16 price-based features from weekly price data and trained four classifiers to predict the direction of the 5-day post-earnings abnormal return relative to the S&P 500. All four used walk-forward validation, so no model ever saw the future during training.
Part 3: Do they agree?

Mostly, no. Sentiment and price-based predictions agreed on only 44.9% of earnings events, below the 50% you'd expect by chance, and the correlation between sentiment and abnormal return was essentially zero (r = −0.043).

That is a meaningful negative result. It is consistent with the semi-strong form of the Efficient Market Hypothesis: by the time information is published in a filing, it is already reflected in the price.
Takeaways
- A careful null result, with walk-forward validation and a clear baseline, is more useful than an overfit positive one.
- Most of the work in applied NLP is in the data: choosing the right sections, chunking for the model's context window, and aggregating chunk-level labels into a defensible document-level signal.