Demand planning is the railway network that turns a destination into a route.

Demand Forecasting Methods: Types, Selection, and What Actually Moves Accuracy

Demand forecasting methods fall into three families. Qualitative methods use structured human judgment. Time-series methods extrapolate from the history of the series itself. Causal and machine learning methods relate demand to outside drivers such as price, promotion and distribution. Which family fits depends on what your data contains: how much history you have, at what level of granularity, and which drivers are recorded alongside it. And in most planning organizations, the accuracy you will reach was set before anyone picked a method.

A planning team is six weeks into an evaluation. Someone has built a comparison grid. Down the left are the method families: exponential smoothing, ARIMA, gradient boosting, a neural network the data science team is enthusiastic about. Across the top are the vendors. The grid is thorough, and by the time it is finished the team will have spent more hours choosing a method than the choice is worth.

This guide gives you the taxonomy anyway, because you need it, along with a selection table you can use on Monday and a way to score methods that does not flatter the winner.

What demand forecasting is

Demand forecasting is the practice of estimating future customer demand for a product at a specific level of detail and over a specific horizon. A demand forecast has four parts: what is being forecast (the item, family or material), where (the location, channel or customer), when (the period and horizon), and how much (the quantity, ideally with a range around it).

Two of those four cause most of the trouble.

Granularity is how finely the forecast is cut. A forecast at product-family and national level is a different object from a forecast at SKU, pack size and distribution center. Aggregate forecasts are almost always more accurate, because errors on individual items cancel each other out. You don’t buy against them. The useful level is the one your purchase orders and production schedules are written at, and accuracy quoted at any coarser level is a flattering number.

Horizon is how far out the forecast runs, and it should be set by your longest committed lead time rather than by the calendar. If a component takes sixteen weeks to land, a four-week forecast tells you nothing you can act on. Most demand plans run 12 to 18 months, reviewed monthly, with the near periods forecast at finer granularity than the far ones.

Three terms then get used as though they were interchangeable.

Term

What it produces

Who owns it

What it is for

Demand forecasting

A quantitative estimate of future demand at a stated level and horizon

Demand planning

An input to every downstream plan

Demand planning

An agreed demand plan, the forecast adjusted for promotions, launches, listings and judgment

Demand planning, with sales and marketing

The number the business commits to operate against

Sales forecasting

Expected bookings or revenue, usually by account and rep

Sales, reported to finance

Quota, pipeline coverage and revenue guidance

The practical difference is what happens when the numbers disagree. A sales forecast that misses by 10% is a revenue conversation. A demand forecast that misses by 10% at SKU and location is expedited freight, a stockout in one region and a write-off in another. Sales forecasts are usually built bottom-up from accounts and are optimistic by construction. Demand forecasts are measured in units, at a level fine enough to buy and build against.

The three families of demand forecasting methods

Family

What it uses

Typical methods

When it fits

Qualitative

Structured human judgment

Delphi, sales force composite, market research, executive judgment

No usable history, or a structural break history cannot describe

Time series

The history of the series itself

Naive, seasonal naive, moving average, exponential smoothing, Holt, Holt-Winters, ARIMA, Croston

Stable patterns, enough cycles to estimate them, no recorded drivers

Causal and machine learning

Demand plus external variables

Regression, gradient boosting, neural approaches, ensembles

Drivers recorded at the same level as demand, and enough series or periods to learn from

Most production forecasts are a combination. A statistical baseline, driver adjustments where drivers exist, and judgment applied on top by a planner. The families describe which part of the forecast is doing the work. Most teams run all three at once.

Qualitative methods

Qualitative methods use structured human judgment instead of, or on top of, historical data. They are treated as the primitive end of the taxonomy. They are also the only option in more situations than a statistics-first reading admits: new products with no history, a market you entered last quarter, a channel that did not exist last year, a tariff regime nobody has operated under before.

Method

How it works

Best used when

Main failure mode

Delphi

Experts forecast independently, see anonymized group results, then revise over two or more rounds

The horizon is long, the question is structural, and the loudest person in the room should not win

Slow. Needs a facilitator and real panel discipline

Sales force composite

Field reps forecast their accounts, and the numbers roll up

Account-level knowledge is genuinely held in the field, and the roll-up covers a manageable number of customers

Optimism, and quota incentives pointing at the number

Market research

Surveys, concept tests, panel data and category read

Launching into a category where you have no analog of your own

Stated intent is a weak predictor of purchase

Executive judgment

Leadership sets the number

Nothing else is available, or the plan is a commitment rather than a prediction

The number becomes a target, and the plan is fitted to it afterwards

A review of the controlled comparisons in the Delphi literature put Delphi panels ahead of statistically aggregated individual forecasts in twelve cases against two, with two drawn, and ahead of conventional face-to-face groups in five cases against one (Rowe and Wright, 1999). The gain comes from the structure: independent estimates, anonymous feedback and revision.

What planners do to a statistical forecast matters too. A study of over 60,000 real forecasts at four companies scored the system’s number against the planner’s revised one. At three of the four, the revisions were a net gain. The pattern underneath that average is the useful part. Substantial revisions usually earned their keep and marginal ones usually did not. Revisions that pushed the number up helped far less often than revisions that pulled it down (Fildes, Goodwin, Lawrence and Nikolopoulos, 2009). Judgment is a real input with a known bias profile, and the fix is to capture it properly.

Time-series methods

Time-series methods extrapolate from the history of the series itself. No external variables, no causal story, just the pattern in the numbers.

To make the comparison concrete, here is one small dataset carried through every method below. Quarterly unit demand for a single SKU over three years, with a mild upward trend and a strong fourth-quarter season.

Quarter

Year 1

Year 2

Year 3

Q1

100

110

124

Q2

130

148

160

Q3

120

132

146

Q4

170

190

208

Every method below is fitted on years 1 and 2, then used to forecast the four quarters of year 3. Year 3 actuals are held back and used only to score.

Naive. The forecast for the next period is the last observed value. Here that is 190 for all four quarters of year 3.

Seasonal naive. The forecast is the value from the same period one cycle ago: 110, 148, 132, 190. One line of arithmetic, and it respects the season.

Moving average. Average the last k periods. With k = 4: (110 + 148 + 132 + 190) / 4 = 145.0, flat across the horizon.

Weighted moving average. Same idea, with more weight on recent periods. Using 0.4, 0.3, 0.2, 0.1 from most recent backwards: (0.4 × 190) + (0.3 × 132) + (0.2 × 148) + (0.1 × 110) = 156.2, flat.

Simple exponential smoothing. Every new observation updates a running level, weighted by a smoothing constant α: new level = (α × latest observation) + ((1 − α) × previous level). With α = 0.3 and the level started at 130 (the year 1 average), the updates run 124.00, 131.20, 131.44, 149.01. The forecast is 149.0, flat. Simple exponential smoothing cannot produce a trend or a season by construction.

Holt’s linear method. Adds a second smoothed component for trend. With α = 0.3 and β = 0.2, starting from level 130 and trend 3.75, the series ends at level 154.58 and trend 6.14, giving 160.7, 166.8, 173.0, 179.1. The forecast now slopes upward and still has no season, so it overshoots Q1 badly and undershoots Q4.

Holt-Winters. Adds a third component for seasonality, multiplicative here. Starting from level 130, trend 0 and seasonal factors of 0.77, 1.00, 0.92 and 1.31, then smoothing with α = 0.3, β = 0.1 and γ = 0.3, the series ends at level 142.5, trend 1.04 and factors of 0.786, 1.021, 0.926 and 1.317. Forecasts: 112.7, 147.6, 134.8, 193.1.

ARIMA. Models the series through its own lags, differencing and moving-average error terms, written as ARIMA(p,d,q) with a seasonal extension. On eight observations it cannot be estimated. A seasonal ARIMA on quarterly data needs several years of clean history before the parameters mean anything, which is the practical constraint on the method, not its mathematics.

Scored against year 3 actuals:

Method

Year 3 forecast (Q1–Q4)

MAPE

Naive

190, 190, 190, 190

27.7%

Moving average, k = 4

145.0 flat

14.3%

Weighted moving average

156.2 flat

15.1%

Simple exponential smoothing, α = 0.3

149.0 flat

14.4%

Holt, α = 0.3, β = 0.2

160.7, 166.8, 173.0, 179.1

16.6%

Seasonal naive

110, 148, 132, 190

9.3%

Holt-Winters, α = 0.3, β = 0.1, γ = 0.3

112.7, 147.6, 134.8, 193.1

7.9%

ARIMA

Not estimable on 8 observations

n/a

Read the ordering, not the decimals. Every method that ignores the season lands between 14% and 28%. Both methods that respect it land under 10%. Seasonal naive is the crudest arithmetic on the list. Holt-Winters beats it by 1.4 points.

Grid-searching the Holt-Winters parameters against those same four held-out quarters gets MAPE down to 0.6%, which describes the holdout and predicts nothing. Tuning to the period you are scoring on is the most common way a pilot produces a number the production system never reproduces.

Intermittent and lumpy demand

The methods above break on intermittent demand. When demand arrives in occasional lumps with zeros in between, smoothing the series averages the zeros into the forecast and systematically under-forecasts the size of an order when one arrives. Croston’s method separates the two problems, smoothing the demand size and the interval between demands separately (Croston, 1972). Croston’s original estimator carries a positive bias, later quantified and corrected in the literature (Syntetos and Boylan, 2001). If a meaningful share of your SKUs are intermittent, get this right before you compare anything else.

Causal and machine learning methods

Causal methods bring in variables outside the series. Machine learning methods do the same with far more flexible functional forms and much larger appetites for data.

Family

What it adds

What it needs before it is worth trying

Regression

An explicit relationship between demand and drivers such as price, promotion depth, distribution points, weather

Driver history at the same level and frequency as demand, and a forecast of the drivers themselves for the horizon

Gradient boosting

Non-linear interactions across many features, learned across many series at once

Thousands of series or many periods, engineered features (calendar, price, promo, hierarchy), and disciplined backtesting

Neural approaches

Shared structure learned across a whole portfolio, useful for long horizons and sparse items

Large volumes of history, tuning capacity, and tolerance for long training cycles

Ensembles and combinations

Averages away the idiosyncratic errors of individual models

Several models already working, plus a rule for weighting them

The competition evidence cuts both ways. In the M4 competition, across 100,000 series, 12 of the 17 most accurate methods were combinations of mostly statistical approaches, and the winner was a hybrid of exponential smoothing and a neural network. The six pure machine learning methods all did poorly. None beat the combination benchmark, and only one beat a seasonally adjusted naive forecast (Makridakis, Spiliotis and Assimakopoulos, 2020). An earlier study by the same authors found statistical methods outperforming machine learning ones on monthly series (Makridakis, Spiliotis and Assimakopoulos, 2018).

Then the M5 competition ran on 42,840 hierarchical series of Walmart unit sales, which is retail demand data with prices, promotions and calendar effects attached. Machine learning won decisively. Every one of the top 50 entries used gradient boosting, and the winner beat the strongest statistical benchmark by 22.4%. About 92.5% of entrants failed to beat that same benchmark (Makridakis, Spiliotis and Assimakopoulos, 2022).

The two results are consistent. M4 series came with no explanatory variables. M5 series came with the drivers attached. Machine learning wins when there is something to learn beyond the shape of the line, and it loses to a well-chosen smoothing model when there is not. The 92.5% figure is the number to carry into a vendor conversation. Nearly every entrant had the same method family available, and most of them lost to exponential smoothing.

Separately, a review gathered 97 formal comparisons of simple against complex methods across 32 papers. The simpler method was as accurate or better in 81 of them, and across the 25 papers reporting quantitative results the added complexity raised error by an average of 27% (Green and Armstrong, 2015).

Hierarchical forecasting and reconciliation

Most planning organizations need the same demand at several levels at once. Finance wants it by category and quarter, procurement by material and month, and the distribution center by SKU and week. Forecast each of those independently and they will not add up, and the mismatch turns into inventory in the wrong building.

There are four ways to reconcile them.

  • Bottom-up. Forecast the most granular level and sum upward. Preserves detail, and inherits all the noise of forecasting sparse items individually.

  • Top-down. Forecast the aggregate, where accuracy is highest, then split it down using historical proportions. Stable at the top, and the split is where the error goes.

  • Middle-out. Forecast at the level with the best signal-to-noise, often product family by region, then aggregate up and disaggregate down.

  • Optimal reconciliation. Forecast every level independently, then adjust all of them so they are coherent, weighting each by how reliable it has proved. It takes more computation and usually beats the other three (Wickramasuriya, Athanasopoulos and Hyndman, 2019).

Where you forecast determines which errors cancel and which compound. It can move accuracy as much as the model family used at that level.

How to judge a method before you commit to it

Method comparisons are only as good as the scoring, and most in-house comparisons are scored in ways that flatter the winner.

MAPE vs. WAPE

Pick the metric that fits your data. MAPE is the default and it breaks where demand is small or intermittent, because dividing by a near-zero actual produces a percentage nobody can interpret. It also penalizes over-forecasting more heavily than under-forecasting, which pushes a team toward models that run light. WAPE, total absolute error divided by total actual demand, holds up at SKU level and is the safer default for item-level review. MAE and RMSE are in units and are fine within an item, misleading across items of different size.

Forecast bias

Track bias separately from error. A forecast can hold a respectable MAPE and still run 8% light every single month. Absolute error tells you how wrong you were. Bias tells you which direction, and direction is what fills or empties a warehouse. Any review that reports only one of the two is missing half the diagnosis.

Rolling-origin backtesting

Back-test on periods the model has never seen, rolling forward. Fit on history, forecast the next period, move the window, repeat. One fixed holdout is how a 0.6% MAPE happens.

Forecast value added (FVA)

Compare against seasonal naive, always. Forecast value added (FVA) scores each step of your process against the simplest reasonable benchmark: the statistical model, the planner override, the consensus adjustment in the demand review. Steps that do not beat seasonal naive are costing money and should be removed.

Measuring accuracy at the right level

Score at the level and horizon you buy at. A model evaluated at national-monthly and deployed at SKU-DC-weekly has not been evaluated.

The selection table

Find the row that describes your situation. Where several rows apply, the most constraining one wins.

If this describes you

Start here

Why

Under 2 years of clean history

Seasonal naive, simple exponential smoothing, or judgment with structure

Nothing more complex is estimable, and pretending otherwise produces confident noise

2 to 3 years of history, clear seasonality

Holt-Winters or seasonal ARIMA

Enough cycles to estimate a seasonal profile, not enough to learn interactions

3+ years, many drivers recorded (price, promo, distribution)

Regression, then gradient boosting

The drivers are the information. A univariate model discards them

Thousands of SKUs, short life cycles

Gradient boosting trained across the portfolio

Cross-learning covers items that individually have too little history

Heavy promotional or trade calendar

Causal model with promotion features, or a baseline plus uplift model

The promotional calendar drives demand more than the season does

Intermittent or lumpy demand

Croston or a bias-corrected variant, judged on inventory outcomes

Smoothing zeros into the level under-forecasts order size

Long horizon (12+ months), structural change

Qualitative methods, scenarios, Delphi

No extrapolation survives a regime change

New product, no history

Analogs, market research, attribute-based models

There is no series to extrapolate.

Small team, hundreds of SKUs or more

Automated model selection with exception-based review

Planner attention is the scarce input. Spend it on the 20% of items that carry the margin

Multi-echelon network, demand at several levels

Hierarchical methods with reconciliation

Independent forecasts at each level will not add up, and the mismatch becomes inventory

What the taxonomy does not tell you

Every method above is a general shape fitted to your history. Exponential smoothing, ARIMA and gradient boosting differ in how flexible the shape is. Each is a template that arrived before your data did, and fitting pulls your history onto it. That has two consequences, and together they explain why forecasts at so many companies seem to hit a ceiling.

Your vendor already chose the shapes. Nobody chose badly here. A fixed model library was state of the art ten years ago, and most advanced planning systems (APS) are still built that way. Your vendor ships a fixed library of model families, identical for you and every other customer, so the achievable accuracy was set before the project started. That is why accuracy so often rises a few points above the incumbent’s and then stops, whatever tuning follows. The ceiling isn’t a tuning failure. It’s where the shapes ran out.

Every method extrapolates history, and history records what you were able to do, not what was possible. When an item was out of stock, the record shows low demand, not lost demand. Estimating true demand from sales in a lost-sales system is a distinct statistical problem precisely because the observations are censored by your own availability (Nahmias, 1994). The same applies to every constraint you imposed. Allocate a scarce product away from a customer and the history shows that customer wanting less. Cap an account because credit was tight, skip a promotion because there was no coverage, ship late because a lane was full, hold a launch because a co-packer slipped, and the record shows reduced demand in every one of those cases. The bullwhip literature documented one version of this decades ago: when supply is rationed, customers inflate their orders, and the order history stops describing real demand (Lee, Padmanabhan and Whang, 1997).

Fit a model to that record and you forecast your constrained self. Buy to that forecast, constrain the business again, and the next fit is tighter still. Growth gets planned out of the business one cycle at a time. Accuracy can improve the whole time. A forecast that reproduces last year’s ceiling scores well against a year in which you hit the ceiling, so the metric confirms the plan instead of testing it. One re-routed order at a manufacturer did precisely this to every cycle that came after it, traced in Why Predicting Isn’t Understanding.

What Omnifold builds instead

Omnifold’s model sits outside all three families on this page. It isn’t a statistical method like Holt-Winters or ARIMA, a supervised learning method like gradient boosting or a neural network, or a large language model writing a forecast. It is a reasoning model of your supply chain, trained with reinforcement learning to reduce forecast error and directed at an economic objective. Accuracy is the mechanism. Revenue, margin and cash are the target. Each customer gets their own model, developed exclusively for their business, and constraints enter it as constraints instead of being absorbed into the demand signal.

The first thing that changes is what the history means. The model can distinguish a decision you made from behavior the market showed, so a stockout window and a lost customer stop looking identical in the record. The second is the target. The forecast can be aimed at the level your business actually operates on, which is not always the finished SKU, as we set out across make-to-stock, make-to-order and engineer-to-order networks in Every Supply Chain Needs a Different Forecast. Scenarios also become cheap enough to run by the hundred, which changes what the planning meeting is for.

The model works from three inputs, which Omnifold calls the Context Gap. History is what happened, including the decisions that constrained fulfillment. Structure is how the business works: network, lead times, capacity, product relationships. Knowledge is what people know is changing before it shows up in the record. Each asks something specific of your data, and it’s worth checking before any vendor conversation. Items 1 to 3 below are History, 4 to 6 are Structure, and 7 is Knowledge.

  1. Order and shipment history at the level you plan and buy at, with customer and channel identifiers intact.

  2. Inventory positions over time, including the windows when an item was unavailable.

  3. Price and promotion history with actual start and end dates, not a monthly flag.

  4. Lead times by supplier, material and lane, as observed rather than as contracted.

  5. Product master data with attributes and hierarchy, so a new item can inherit from something.

  6. A record of the constraints you imposed: allocations, minimum order quantities, capacity caps, credit holds.

  7. What your planners and commercial teams know that the record doesn’t show yet: a customer’s plans, a supplier’s problems, a launch slipping, a growth target.

Knowledge is where the Fildes findings pay off. Substantial planner revisions usually improve the forecast, and in traditional planning that knowledge is typed over the number as an override and gone by the next cycle. In Omnifold, planners supply it as Enhancements and the model learns from them. The platform proposes Enhancements too, and the planner decides which ones stand.

At Carbliss, a beverage brand and Omnifold customer, forecast error at SKU and distribution center fell 40% against their incumbent, which was itself machine-learning based. The better forecast let Carbliss cut inventory 20% while maintaining service levels, releasing about $4 million in cash.

Omnifold does not fix master data that disagrees with itself, a demand signal nobody records anywhere, a sales team whose numbers are quota rather than prediction, or a business that will not act on the plan it agreed to. It will not make a genuinely unforecastable event forecastable. What it removes is the ceiling that comes from fitting a general shape to a constrained record.

Four signs your method has hit its ceiling

  1. Error is flat despite tuning. Two quarters of parameter work, feature engineering and model selection have moved MAPE by a point or two. That is the library’s ceiling, and more configuration won’t lift it.

  2. Planners override routinely. When overrides are the norm rather than the exception, the model is not carrying the information the business runs on. The Fildes findings say those overrides are often right. The problem is the model, and asking planners to override less won’t fix it.

  3. Error concentrates in the SKUs that carry the margin. Aggregate accuracy looks respectable because the tail is easy. The items with promotions, allocations and constrained history are the ones that miss, and they are the ones that pay for everything.

  4. New products are wrong every time. A model with no mechanism for items without history will be wrong on every launch, and launches are where growth lives.

If two or more of these describe your situation, skip the next model bake-off. Check what your history is actually recording, and whether your forecast can tell your decisions apart from your market.

Frequently asked questions

What are the main types of demand forecasting methods?

Three families. Qualitative methods use structured judgment (Delphi, sales force composite, market research, executive judgment). Time-series methods extrapolate the history of the series itself (naive, moving averages, exponential smoothing including Holt and Holt-Winters, ARIMA, Croston for intermittent demand). Causal and machine learning methods relate demand to external drivers (regression, gradient boosting, neural approaches, ensembles).

Which demand forecasting method is most accurate?

It depends on what your data contains rather than on the method’s sophistication. In the M5 competition, where series came with price, promotion and calendar data attached, gradient boosting methods swept the leaderboard and the winner beat the best statistical benchmark by 22.4%. In M4, where series came with no explanatory variables, the winners were combinations, led by a hybrid of exponential smoothing and a neural network, and the pure machine learning entries finished near the bottom. Match the method to the information available, and expect a well-chosen simple method to get close to a complex one.

How much history do I need for demand forecasting?

Two full seasonal cycles is the practical minimum for estimating a seasonal profile, and three or more before a seasonal ARIMA or a driver-based model is worth the effort. Below that, seasonal naive or simple exponential smoothing with structured judgment will outperform anything more elaborate.

What method should I use for intermittent or lumpy demand?

Croston’s method or a bias-corrected variant, and judge it on inventory outcomes rather than on percentage error, because percentage error metrics behave badly when actuals are frequently zero.

What is the difference between demand forecasting and demand planning?

Demand forecasting produces a quantitative estimate from data and models. Demand planning is the process that turns that estimate into a plan the business commits to, incorporating promotions, launches, listings and judgment. The forecast is an input. The demand plan is a decision.

Which forecast accuracy metric should I use?

WAPE for item-level review, because it holds up where demand is small or intermittent and MAPE does not. Track bias alongside it, since a model can post an acceptable error rate while running consistently light or heavy. Compare every step of your process against a seasonal naive benchmark and remove the steps that fail to beat it.

Does AI actually improve demand forecasting?

It does where there is something beyond the shape of the series to learn from, which is most consumer-facing supply chains. It does not where the data is thin or the drivers are unrecorded. The larger gains usually come from what the model is built on rather than from the algorithm family.

Can we just use the method our ERP or planning system ships with?

You can, and it will give you a defensible baseline. Understand that the system is selecting from a fixed library of general model shapes, which sets a ceiling on how accurate it can become for your business, and that the ceiling is mostly a function of the fit between that library and your supply chain rather than of your configuration effort.

Related reading

Sources