How people trick themselves into thinking they can predict the stock market
I keep accumulating ideas of things to make into YouTube videos or long careful tutorials or whatever else. And then I never do them. So the new plan is to just do quick blog posts when I think about something. If one day I come back around and cover the same material in more detail then good! Otherwise at least it’s out there.
There is a whole genre of YouTube videos showing how to predict the stock market. Lots of these videos use relatively complex models, namely lstm neutral networks. Most of these videos use very simple data, namely the price of the same stock at previous times. These videos typically conclude (and splash loudly on the thumbnail) that they can predict the stock market with appreciable accuracy.
A few simple arguments, that have been well made elsewhere, indicates that they are probably wrong. If they can predict the stock market, they will be off making millions, not making YouTube videos. If they can easily predict the stock market, then so can everyone else. If everyone can predict the stock market, any predictable signal from the data will become priced in and the whole process will quickly fall apart.
So the question is, why does these people’s analyses say they can predict the stock market, when we can be pretty sure that they, in fact, can’t. Like many things, the ‘why’ is more interesting than the fact itself.
I think there’s at least four reasons, some of which are quite subtle. These reasons are interesting both directly for people interested in the stock market, but also interesting for anyone interested in forecasting or predictive modeling more generally. For the benefit of readers with short attention spans, I’ll start with the most interesting, subtle reason, that I haven’t seen discussed at length before. I’ll then circle back to the more obvious answers.
Data leakage from stock choice
If I asked you to tell me everything you know about Tesla (especially if you are interested enough in the stock market to be reading this post) you might answer something like “electric cars, Elon Musk, massive stock price growth”. Notably, most of the videos trying to predict the stock price use one of Tesla, Apple or an index fund like spy. With all of these stocks, we have (probably unconciously) used our knowledge of the present to select them. These are stocks that have, on average, gone up.
This simple fact means the prices of these stocks suddenly are predictable to an extent. A model that predicts a small, positive change in the stock price will do better than random. However, stocks go up, until they stop going up. If you took a random stock, or many random stocks, and tried to predict the price in the future we wouldn’t have this small guarenteed predictive ability. Similarly, if you took the model that predicts a small positive change in the stock price and applied it, long term, to Apple or Tesla, your predictive accuracy would depend entirely on whether these stocks keep going up or not. Eventually they’ll come down, all companies eventually go busy.
This point is quite subtle and I haven’t seen people make it before. But it applies beyond the stock market. For example, if you predict something about species population size or epidemic size, but only use species or diseases that haven’t gone extinct, your models will perform better than they would in the real world.
Predicting value, not change
I think the actual biggest reason that people on youtube trick themselves into thinking they can predict the stock market is that they often try to predict stock price rather than the change in the stock price. Stock prices change through time, but generally not hugely. Other ways of saying this is that the stock price today depends a lot on the stock price yesterday, or that the stock price is an autoregressive process.
So for a stock that has changed price a lot over a long period, such as Tesla that is now worth much more than it was 10 years ago, a model that predicts tomorrows price as being the same as todays price, will have very high apparent “predictive ability”. When the price went from $1 to $1.1, you predict $1. When the price went from $100 to $110, you predict $100. The correlation between your predictions ($1 and $100) and the truth ($1.1 and $110) is high.
However, the problem with this is that any benefit in predicting the market gained by this property, is exactly cancelled out by the fact that to make money off a stock today, you have to have bought it yesterday. The value of a stock doesn’t have any bearing on your profit, only the change in the value of the stock.
Using graphs as metric
A related problem is that many videos fit models, make predictions of the stock price then plot the predictions against the truth in a typical time-series line plot. The lines on these plots often follow each other and look quite convincing. However, this approach fails in the same way as the issue of predicting value, not change. A model that predicts that the stock price tomorrow is the same as the stock price today will look pretty good on these plots. But they are unfortunately utterly useless in terms of making money on the stock market.
Using data from the future
The final, least subtle point, is that it’s easy to accidentally use data from the future. I think most of the youtube videos are actually quite careful on this point, but I thought I’d include it for completeness.
It is off course obvious that if you use tomorrows stock price as a predictor in your model, you will be able to predict tomorrow’s stock price! However, there is a risk of accidentally using more subtle information from the future. If you are using other stocks as predictors, you need to make sure you are using todays stock price not tomorrows. Imagine you are predicting the change in price of Pepsi stocks, but using changes in Coca-cola stock prices as a predictor. If the government announces a sugar tax, both stock prices will fall (I guess). If you accidentally use tomorrows change in Coca-cola stock price, you are accidentally telling your model that there will a sugar tax announced tomorrow, but this is information you would not have in a real model. Relatedly, in the process of calculating compound variables such as moving averages, open-close, high-low you can accidentally use future information. A model that buys when the price hits the weekly high is using information from the future as you don’t know what the high is until the end of the week, at which point you’ve missed the high you were hoping to buy at.
Conceptually this is mostly quite simple. The problem is that it is easy to mess it up in your code if you are not careful.
Final thoughts
So overall, it’s actually essentially impossible to predict the stock market, in a useful way, using stock prices, desktops and tens of minutes of effort. You are always going to be slower than the high-frequency trading algorithms, and anything simple will have already been done. It’s still fun to try, but it’s mostly a humbling experience that tests your ability to implement everything carefully. It’s also a fun exercise in knowing when to give up, when to conclude that the variable of interest is impossible to predict given the available data. This is a topic I’d like to write more about.
Interestingly, if I remember correectly, predicting the change in volume (the number of shares bought and sold) is possible. Unfortunately, it’s not easy to know how to make money with that information. But it’s still perhaps a fun game to play.
A practical guide to squishing yearly ‘disag data’ objects
A practical guide to pooling annual data for Bayesian disaggregation regression in R
This post is a guest post by my predoctoral fellow, Samana.
Bayesian spatial disaggregation gives you a high resolution risk surface from low-resolution case counts ….. but only one year at a time. This post shows how to pool several years into a single model fit.
Hi all, my name is Samana and I’m a predoctoral fellow and I’ve been working on the problem of a multi-year disaggregation model!
In this blog I will be presenting a recipe of how to do a multi-year disaggregation model fit.
Currently, using the disaggregation package we can only really fit one model per time slice of data we might have, but this blog aims to change that and let you fit a disag model with multiple years of data.
Part A demonstrates the mechanics without a time trend; Part B adds a time spline to the six-year model. This blog is part A.
This post shows how to combine annual disag_data objects and fit one pooled spatial disaggregation model. It demonstrates shared effects across years and it is not a full spatio-temporal model with separate spatial field for each year.
The resolution mismatch problem:
Simply put, there is a resolution mismatch between satellite covariates (which are very high resolution) and public health data, like case counts of a particular disease in a specific county (this is low resolution).
Public health data is often reported as annual case counts for administrative areas such as counties and districts whereas environmental information that may explain risk, such as temperature, land cover, vegetation is often available on a fine spatial grid.
The problem is that we can’t just combine these together and then fit a model for risk predictions to get a high resolution risk surface. This is because of the ecological fallacy.
These figures above show the resolution mismatch; on the left we have the low resolution public health case counts and on the right we have the high resolution environmental satellite covariates (Nandi et al. 2023, https://doi.org/10.18637/jss.v106.i11).
What is disaggregation and why you’d want to do it:
So these resolutions are linked by modelling a latent disease risk at the pixel level and THEN aggregating the pixel level predictions to the polygon level so they’re consistent with the observed disease counts. all of this is done by Bayesian spatial disaggregation (implemented in the R package disaggregation )
The package fits a Bayesian disaggregation regression model (similar to a GLM) with an INLA style spatial random field via TMB, using the function prepare_data() to line up polygons, covariates and population, and disag_model() to fit the model.
The current single year workflow is like this:
Read in the inputs: a sf object containing the area boundaries and case counts, a SpatRaster of fine-scale covariates, and a population raster when modelling counts with population as exposure.
Prepare these inputs into a coherent data object with prepare_data().
Fit the model with disag_model(), choosing the likelihood and link and whether to include a spatial field and an independent (iid) effect.
Inspect the fit with summary() and plot(), then use predict() to produce fine-scale prediction and uncertainty maps.
For multiple years, the missing step is between 2 and 3: prepare each year separately, then combine the resulting disag_data objects carefully before fitting one pooled model.
Why multi-year disaggregation?
When you have many years of polygon level counts, a natural question is: fit each year separately, or pool several years into one model?
Pooling helps when you want the spatial field and covariate effects to borrow strength across years, while still letting a time trend move through a flexible term (we’ve used a natural spline on centered year in part B of this blog).
I’ve been calling this “squishing” years together, stacking each year’s prepared data into one long object before handing it to disag_model().
Bayesian spatial disaggregation is straight forward for one year of data but multi-year models are harder because the disag_data object contains many components that need to be in the right shape, format and size.
This object contains polygon-level outcomes, pixel-level covariates, aggregation weights, spatial coordinates, start end indices and the spatial mesh. These objects must remain correctly aligned when the years are combined.
This tutorial develops a reproducible workflow for combining yearly disag_data objects informally, “squishing” or “combining” them before fitting a multi-year Bayesian disaggregation model.
Model components
The disaggregation regression models demonstrated here include fixed effects of covariates and a number of random effects, for various different modelling purposes.
In this example, the pooled model estimates shared climate effects. We have different covariate observations for each year, but want to fit a model using all years of data in one go.
The spatial random effect operates at high resolution. If you model each year separately, then each time period has its own spatial field, defined over the mesh and projected to the pixels. It captures smooth residual spatial variation that remains after accounting for measured environmental covariates. However, we often want to pool estimates for this random effect across years.
The iid random effect operates at polygon level, the lower resolution. It gives each polygon-time observation an additional independent residual term, helping account for overdispersion and unmeasured area-level factors.
When yearly data are squished, the model is fitted jointly across all periods. The workflow must preserve the relationships between each polygon, its pixels, its aggregation weights and its appropriate spatial field.
The iid term allows an additional residual effect for each polygon-year observation. Stacking the data does not automatically create a separate spatial field for each year or model how fields evolve over time.
THE DATA:
The example region is Madagascar, deliberately chosen because it’s the same country used in the disaggregation package’s own founding methods paper (Nandi et al. 2020, malaria case counts across Madagascar, [Nandi et al](https://www.jstatsoft.org/article/view/v106i11)) so the mesh settings and covariate naming are the same as that original worked example.
The geometry, climate covariates AND population is real public data; only the case counts are simulated (real disease data can’t be shared here for data-permission reasons)
The dataset pairs real, public district geometry, real climate covariates and real population with simulated case counts:
Real public administrative level-2 (district) boundaries for Madagascar, from the GADM database (https://gadm.org), downloaded via the R geodata package.names.
Year varying covariates (temp_wc, precip_wc): annual mean temperature and annual total precipitation for each simulated year, built from REAL monthly TerraClimate records (Abatzoglou et al. 2018) not long-run climate normals, so these differ from one year to the next. Centred and scaled (mean 0, sd 1, using one mean/sd computed across all years combined so real year-to-year differences are preserved). Source: https://www.climatologylab.org/terraclimate.html.
Simulated case counts drawn from known truth parameters (shipped alongside the data as true_parameters.csv), using the real covariates and population above plus a simulated spatial random field and district-level noise. So this post can check whether the fitted model recovers the true coefficients, a much better teaching device than fitting to arbitrary noise.
None of the numbers in the dataset correspond to any real disease case count.
:::
### Step 4: prepare data for each year
Mesh args here match those used in the `disaggregation` package's own
Madagascar malaria worked example (Nandi et al. 2020) a nice side
benefit of picking the same country is that these don't need retuning
from scratch.
Always retune these to your own study area if you swap in a different
region.
The details for what these parameters mean can be found in the inla
tutorials for example.
::::::: cell
``` {.r .cell-code}
mesh_args_mdg <- list(max.edge = c(0.7, 8), cutoff = 0.05, offset = c(1, 2))
dis_2001 <- prepare_data(
polygon_shapefile = shapes_2001,
covariate_rasters = cov_stack_2001,
aggregation_raster = population_raster,
mesh_args = mesh_args_mdg,
id_var = "district_id",
response_var = "cases",
na_action = TRUE
)
::: {.cell-output .cell-output-stderr} Warning: [rast] CRS do not match :::
::: {.cell-output .cell-output-stderr} Warning: [extract] transforming vector data to the CRS of the raster :::
::: {.cell-output .cell-output-stderr}
Warning: [rast] CRS do not match
Warning: [extract] transforming vector data to the CRS of the raster
:::
``` {.r .cell-code}
dis_2003 <- prepare_data(
polygon_shapefile = shapes_2003,
covariate_rasters = cov_stack_2003,
aggregation_raster = population_raster,
mesh_args = mesh_args_mdg,
id_var = "district_id",
response_var = "cases",
na_action = TRUE
)
::: {.cell-output .cell-output-stderr} Warning: [rast] CRS do not match Warning: [extract] transforming vector data to the CRS of the raster ::: :::::::
Step 5: The squish: stack the 3 years of data into one disag data object
To stack these objects together correctly, we handle different parts of the data differently. polygon_data and covariate_data are just stacked with a year label attached. We offset start_end_index: each prepare_data() call starts its pixel-row numbering again, so the second and third years’ indices must be shifted to point to their new positions in the combined covariate table.
::: cell ``` {.r .cell-code}
polygon_data and covariate_data just stack with a year label attached
poly_0103 <- bind_rows( transform(dis_2001$polygon_data, year = 2001), transform(dis_2002$polygon_data, year = 2002), transform(dis_2003$polygon_data, year = 2003) )
cov_0103 <- bind_rows( transform(dis_2001$covariate_data, year = 2001), transform(dis_2002$covariate_data, year = 2002), transform(dis_2003$covariate_data, year = 2003) )
:::
### step 6: fit the pooled model
:::::: cell
``` {.r .cell-code}
fit_0103 <- disaggregation::disag_model(
data = dis_0103,
iterations = 1000,
field = TRUE,
iid = TRUE,
family = "poisson",
link = "log"
)
::: {.cell-output .cell-output-stderr} Fitting model. This may be slow. :::
Compare that against three separate single-year fits (disag_model() on 2001, 2002, 2003 individually, no squishing) to see whether pooling changes the covariate estimates, tightens their standard errors, or shifts the spatial field. This comparison is the actual point of the exercise the code above just gets you to a model you can compare against.
:::::: cell ``` {.r .cell-code} fit_2001 <- disaggregation::disag_model(dis_2001, iterations = 1000, field = TRUE, iid = TRUE, family = “poisson”, link = “log”)
::: {.cell-output .cell-output-stderr}
Fitting model. This may be slow.
:::
``` {.r .cell-code}
fit_2002 <- disaggregation::disag_model(dis_2002, iterations = 1000, field = TRUE,
iid = TRUE, family = "poisson", link = "log")
::: {.cell-output .cell-output-stderr} Fitting model. This may be slow. :::
``` {.r .cell-code} fit_2003 <- disaggregation::disag_model(dis_2003, iterations = 1000, field = TRUE, iid = TRUE, family = “poisson”, link = “log”)
::: {.cell-output .cell-output-stderr}
Fitting model. This may be slow.
:::
::::::
::::::: cell
``` {.r .cell-code}
summary(fit_2001)
Pooling the three years reduced uncertainty: the pooled model had smaller standard errors for both temperature and precipitation than any separate annual model. The temperature estimate was positive and the precipitation estimate negative, consistent with the values used to simulate the data.
Refrences
Nandi, A., Lucas, T., Arambepola, R., Python, A. (2023). disaggregation: An R Package for Bayesian Spatial Disaggregation Modeling. Journal of Statistical Software, 106(11). https://doi.org/10.18637/jss.v106.i11
Nandi, A., Lucas, T., Arambepola, R., Gething, P., Weiss, D. (2020). disaggregation: An R Package for Bayesian Spatial Disaggregation Modelling (original methods paper, Madagascar malaria worked example). https://arxiv.org/abs/2001.04847
GADM database of global administrative boundaries: https://gadm.org
Abatzoglou, J.T., Dobrowski, S.Z., Parks, S.A., Hegewisch, K.C. (2018). TerraClimate, a high-resolution global dataset of monthly climate and climatic water balance from 1958-2015. Scientific Data, 5, 170191. https://doi.org/10.1038/sdata.2017.191