Back to Blog

LLM narratives and macroeconomic data: the Murex research proposal Émile shared

2 min read
“Murex’s 2025 LLM Macroeconomic Research Proposal” sits beside a diagram of networks, charts and a computer labelled “Applied Mathematics”, “Time Series Analysis”, “Bayesian Learning” and “Natural Language Processing NLP”.

On 6 November 2025, Émile Servan-Schreiber reshared a doctoral research invitation from Charles-Albert Lehalle with a wistful comment about being younger. The proposed project at Murex would use large language models to extract time series from narratives, then combine those series with market and macroeconomic data. It was a concrete research direction, rather than a claim that language models had already improved macroeconomic forecasts.

Turning narratives into a measurable input

Lehalle described work in the Paris area, in collaboration with Anna Simoni and himself. His announcement named three data frequencies: quarterly, weekly and daily. Combining them creates a practical timing question: a narrative may change today while an official economic series updates much less frequently. The proposed research would bring those different inputs into macroeconomic scenario generation.

The proposed division of labour is useful to understand. A language model would help represent changing narratives as a sequence of observations. Those observations would then sit alongside numerical data in the scenario-generation process. Reading economic commentary and forecasting the economy are separate steps: the ability to extract a narrative does not demonstrate that it contains information about what happens next.

A research precedent, with an important boundary

Simoni and Laurent Ferrara’s earlier paper, When are Google data useful to nowcast GDP?, studies alternative information from online searches. Their method combines variable selection and ridge regularization, which constrains fitted coefficients, and examines performance outside the data used to fit the model. The authors report that search data can add information beyond official variables, with gains differing between recessions and stable periods.

That paper concerns Google searches and GDP nowcasting: estimating current economic activity before complete official figures arrive. Narrative monitoring introduces a different source of information. A useful evaluation would ask whether the extracted series add predictive information beyond market and official data, and whether that contribution holds outside the period used to build the model. This is a methodological question raised by the proposal, not a reported result of the doctorate.

For readers following the journal’s work on language models and forecasting accuracy, the relevance is the connection between language-derived information and a quantitative forecasting process. Servan-Schreiber’s reshare expressed interest in that research direction. The announcement is evidence of a proposal; its accuracy would have to be established through subsequent testing.

Evidence

Sources & further reading

Keep reading

Related articles

InsightsCrowds versus language models: the more accurate the AI forecaster, the more it errs like humansInsightsAI Forecasting vs Prediction Markets: How to Compare the SignalsInsightsHugo, Galton and the clues a crowd can put together

Subscribe to The Forecasting Brief — forecasts & model notes, once a week