Skip to content

About Automatic Statistician

A publication about automated statistical analysis, interpretable models and the machinery that turns raw data into an explanation a person can read, check and disagree with.

Mission

Make sophisticated data analysis easier to understand, automate and apply.

Most analytical work still fails for unglamorous reasons. The wrong model is chosen and never questioned. Seasonality is mistaken for trend. A structural break is averaged away. A confident number is reported without the interval that should sit beside it. Automation can remove a great deal of that drudgery, but only if the output remains something a human being can audit.

That is the line this site works along: methods that automate the search for structure while keeping the result legible. We publish explanations of the underlying statistics, worked examples on real datasets, and assessments of the software that claims to do this for you.

What we cover today

  • Automated statistical analysis — model search, data profiling, generated reports
  • AutoML — selection, tuning, evaluation and the limits of each
  • Explainability — interpretation, auditability and model criticism
  • AI analytics — assistive systems for analysts
  • Time series — trend, seasonality, changepoints and forecasting
  • Probabilistic modelling — Bayesian inference and uncertainty
  • Analytical software — reviews measured against a published methodology
  • Practical tutorials — applied guides with runnable examples

Project history

The Automatic Statistician began as a research project asking a deceptively simple question: how much of what a statistician does when they first meet a dataset could be carried out by a machine, without the answer turning into a black box?

The approach was to define an open-ended language of models — Gaussian process kernels composed by addition and multiplication — and search it. Smooth trends, periodic components, linear growth, changepoints and noise could be combined into structures rich enough to describe real series. Candidate models were compared with Bayesian criteria that trade fit against complexity, and the winning structure was then translated into charts and natural-language prose, together with model criticism identifying where the model and the data disagreed. The output was a ten-to-fifteen page report.

The original system was developed at the University of Cambridge by James Robert Lloyd, David Duvenaud and Zoubin Ghahramani, in collaboration with Roger Grosse and Joshua B. Tenenbaum at MIT. Later work in the Cambridge Machine Learning Group extended the project through student and doctoral research, including theses by Nikola Mrkšić, Riaz Moola and Qiurui Charles He. The Research page lists the publications in full.

Those reports are the reason this domain accumulated citations from research groups, universities and the technical press over the following decade. The example analyses of airline passenger numbers, solar irradiance and unemployment data remain the clearest demonstrations of the idea, and they remain available here.

Project History & Stewardship

The Automatic Statistician originated as a research project exploring artificial intelligence for automated, interpretable data analysis. AutomaticStatistician.com is now independently maintained and continues that broader educational and technical mission. Historical research is credited to its original authors and institutions. The present website is not an official University of Cambridge, MIT or Google website.

Who is behind the project today

AutomaticStatistician.com is edited and maintained by Elizabeth Sramek, a B2B strategy advisor and systems architect based in Prague. She has worked on the technical foundations of data-driven businesses since 2005 — measurement, attribution, data infrastructure, and the analytical layer that sits on top of them. She has no association or affiliation with the original researchers, their institutions, or the funders of the historical project.

Her work has consistently circled the same problem from different directions: an organisation collects far more data than it can interpret, and the interpretation is where the value sits. Twenty years of building and optimising commercial systems is, in practice, twenty years of asking what a number means, whether the change in it is real, and what should be done differently as a result. Those are statistical questions long before they are business questions, and they are answered badly far more often than anyone admits.

That is what drew her to this domain. The Automatic Statistician was, and remains, one of the few serious attempts to automate analysis without surrendering the explanation. Most of what is now sold as AI-driven analytics does the opposite: it produces an answer with great confidence and no account of how it got there. A system that can say this series has a twelve-month cycle whose amplitude is growing, and here is the interval around the forecast, and here is where the model disagrees with the data is worth considerably more than one that simply returns a number.

Her current work sits at the intersection of machine intelligence and information architecture — designing systems that are legible to both people and machines, and building the structured signals that let automated systems verify and cite a source rather than paraphrase it. The editorial approach here follows from that: define terms precisely, show the working, publish the dataset and the method alongside the conclusion, and state the limitations in the same breath as the result.

The plan for this site is neither a museum nor a rebrand. The historical pages, example reports and publication records stay where they are and remain properly attributed. Around them, new material covers what the field has become since: AutoML and its failure modes, explainability as an engineering requirement rather than a compliance checkbox, automated exploratory analysis, forecasting, and the software that now claims to do all of it. Where a tool is reviewed, it will have been used, on a stated dataset, against stated criteria.

The longer-term ambition returns to the original promise — upload data, receive an analysis you can actually read — approached carefully, and without claiming capabilities the software does not have.

How we work

Attribution first. Historical research is credited to its authors, with papers, venues and links. New commentary is clearly marked as ours.

Show the working. Examples state the dataset, the question, the method and the limitations. Charts are reproducible from public data.

No invented numbers. Benchmarks and statistics are sourced or measured. Where a tool has not been tested, we say so rather than imply otherwise.

Commercial transparency. Any commercial relationship is disclosed on the page it affects. Research sections carry no commercial placements.

The full methodology and editorial policy set this out in detail.