Data analysis & writing by Milcah M. Joseph
MSc Business Analytics · Dublin Business School
ESG is now one of the loudest signals in a company's public image. So does that attention actually predict how its stock moves?
The public attention around a company, the searches people run, including searches about its sustainability, carries real information about where its stock is headed. But that information has a shelf life of about a week, and then it's gone. A company's ESG rating, it turns out, has nothing to do with whether the signal shows up at all. And the most useful conclusion is the one that sounds like a letdown: the signal is real, but too fleeting and too expensive to keep chasing. What this project is really about is the discipline it takes to prove that honestly. To tell a genuine pattern apart from tens of thousands of lucky accidents, and to know the difference between a result that impresses and a result that is true.
ESG has become one of the loudest signals in modern markets. A shorthand for a company's reputation, its risk, its resilience. Which raises a fair question: if the world is paying this much attention to how companies behave, does that attention actually surface in their share prices? And does a company's sustainability standing change how readable its fortunes are?
Those were the two things I set out to test: 1. whether the online attention around a firm, measured through Google search interest (including interest in its sustainability), improves a forecast of its weekly returns, and 2. whether companies of different ESG ratings behave any differently. I used weekly stock returns as a practical stand-in for near-term financial performance, and Google Trends search volume as a measurable proxy for public attention. Simple questions. The answers took a lot of machinery to earn.
There was no clean dataset waiting for this, so I built one. It came to 45 companies across the US, UK, and India, sorted into high, mid, and low ESG using FTSE Russell ratings pulled across ten years (so that one unusual year could not define a company's label). Against them, 124 search terms across 6 regions, from 2021 to 2024.
The worst of it was hiding in the search data itself. Google Trends does not hand you raw search counts. It scores each batch of terms from 0 to 100, where 100 is simply the busiest point inside that particular batch. Pull two terms in separate batches and both will touch 100 at their own peaks, even if one is searched a thousand times more often than the other. Pulled the obvious way, the numbers look perfectly comparable and are quietly meaningless. So into every single pull, hundreds of them, I wedged one fixed reference term, a constant yardstick that let me rescale every batch onto the same footing. Miss that step, and every comparison in the entire study would have been built on sand.
The rest was the ordinary grind of making imperfect data behave, each fix a small judgment call. Trading weeks end on a Friday, except when a holiday shifts them, while Google's search weeks run Sunday to Saturday. So I realigned every trading week back to its Sunday to make the two calendars line up. I worked in returns, that is, percentage moves, rather than raw share prices. So a company trading at 500 and one trading at 5 sit on the same scale. The models need each data series to be statistically stable over time rather than drifting, so I tested all 785 of them and corrected the one that was not. Search terms that barely moved got dropped, since a flat line carries no information. Coca-Cola trades in both the US and UK, so I kept both listings but counted its name searches only once, so it was not double-weighted. And I threw out 2020 entirely: the pandemic crash was a one-off break in how both markets and search behaved, and training on it would teach the model a world that no longer exists.
None of this is glamorous. It is also most of the actual work, and getting it wrong quietly poisons everything downstream.
First, a wide net. I tested every company against every search term in every region, more than 128,000 combinations in all. But there is a trap in testing that much, and it is the one most attention-and-markets studies fall into: run enough tests and pure chance will hand you significant-looking results even when nothing real is there. Flip enough coins and some will land heads ten times in a row. So before I trusted a single result, I applied a correction that raises the bar in step with how many tests were run (a false-discovery-rate control), which throws out the lucky flukes.
Then, the eye of a needle. Passing a statistics test is a surprisingly small feat, and a result can be statistically significant while being worthless to anyone actually trying to use it. So every surviving signal had to clear a second set of hurdles, the kind a trader would care about rather than a statistician. It had to beat a plain model that looks only at a stock's own past (a stock-only AR benchmark). It had to beat an even plainer one that just assumes next week looks like this week (a naïve “no-change” forecast), and beat it by a clear margin (>10%). And its mistakes had to stay small relative to how much the stock naturally bounces around week to week (forecast error under half the stock's own volatility). Only what cleared all of that counted.
I also added one rule most studies skip. A predictor got credit for working three weeks out only if it had already worked at one week and at two. No going back to cherry-pick a lucky later week after an early miss. Fail early, and it was done.
There is real signal in here. At one week out, search attention improved the forecast for 43 of the 45 companies, and about a fifth of all the candidate models cleared every one of those economic hurdles. Set against how rare survivors were, that fifth might sound underwhelming. It is the opposite. With signals that survive a hard set of gateways, this is exactly the fingerprint of something real hiding in the noise.
And then it's gone. This is where the title comes from. Of the 592 signals that worked one week out, only 158 still worked at two weeks, 48 at three, and 13 at four. Search attention gets absorbed into the price within a single trading week, and after that the trail goes cold.
The edge does not fade gently. It falls off a cliff.
Faint bars mark the week-one baseline. Of 592 signals that work at week one, only 13 hold in four weeks.
And the ESG question, the one I started with? A company's sustainability rating had nothing to do with any of it. When I tested whether high-, mid-, or low-ESG firms were more or less predictable, the result came back flat (statistically: chi-square (2) = 1.84, p = .40). Green companies, brown companies, and everything in between showed roughly the same hit rate. Whatever makes a company's fortunes briefly readable from public attention, its ESG standing is not it.
It would be easy to squeeze more out of this than it can bear, so here is the line I'm drawing. The one-week signal, I am confident in: it is real, it survives the strict tests, and it decays in the same manner across companies and time periods. Almost everything else, I am treating as an observation, not a conclusion. Most of the apparent differences - by region, by company size, by ESG - dissolve once the economic hurdles are applied, and the few that linger are not strong enough to trust at this sample size. Searches from the Worldwide region, for instance, looked over-represented among the winners, until I ran the numbers and found the pattern sat well within what chance would produce (a goodness-of-fit test put it at p ≈ 0.21). With 45 companies, that is a hint to investigate, not a finding to report. Knowing which is which is the whole job.
So is this a usable forecasting tool? No. And saying so plainly is the point. The signal is real, but it lives about a week, and keeping it current would mean re-screening thousands of search terms and refitting every company every single week, an enormous amount of work for a very short look ahead. The deeper limit is the data itself: Google's weekly search figures are coarse and cannot be pulled in bulk, which caps how sophisticated a model can get. And at the same time, the simple, transparent model I used cannot capture the more tangled patterns that might survive longer. A better answer probably exists. It would need richer data and a cleverer model than the current tools allow.
What this project actually delivers is not a forecast. It is the method, and the honesty. A repeatable way to sift a huge, noisy pile of signals under the constant risk of fooling yourself, and to come out the other side able to say which patterns are real, which are luck, and, hardest of all, which are real but still not worth acting on. Most analyses are built to find a win. This one was built to find out what was true, and to say so even when the answer was that it does not work.
First Class Honours (81%), MSc Business Analytics, Dublin Business School.
Built in Python (statsmodels, scikit-learn, pandas, numpy) on data from FactSet and Google Trends.