
AI Demand Planning: What It Changes and What It Still Depends On
Sit through two demand planning demos this quarter and both will promise AI. In the first, it turns out to mean a forecasting algorithm that has been on the market for a decade repackaged as AI. In the second, it means a chat window sitting on top of the same forecast.
Most of what is sold as AI demand planning predicts from your history, faster. How much better your plans get depends on something else: how much of your business the model actually understands. That means your history, the structure of your network and what your planners know. This guide explains the technologies behind the label, what independent research says about them and how to tell which ones will change the decisions your team makes.
What is AI demand planning?
AI demand planning is the use of artificial intelligence to forecast demand and to decide what to make, buy, stock and promise in response. The term covers two jobs that often get blurred. AI demand forecasting produces the prediction: expected demand by SKU, location, customer, channel and period. AI demand planning uses that prediction, together with inventory, capacity and commercial knowledge, to choose a plan.
A forecast can be accurate and the plan built on it can still be poor, so this guide spends most of its time on the second job.
What does “AI” mean in demand planning today?
The AI label in planning covers technologies that work in very different ways. Some predict the next value from patterns in past data, and others work through a problem toward a goal.
Forecasting methods that predict from history (sometimes called time-series analysis). Much of what is marketed as AI demand planning is one of these:
Statistical forecasting. Formulas such as exponential smoothing and ARIMA that extend a product’s past trend and seasonality forward, the backbone of advanced planning system (APS) forecasting engines for decades.
Machine learning forecasting. Algorithms such as gradient-boosted decision trees and neural networks that learn demand patterns from sales history plus drivers like price, promotions and weather. They can learn across thousands of related products at once, which statistical formulas fitted one product at a time cannot.
Demand sensing and exception detection. Two common uses of these methods. Demand sensing corrects the next few days or weeks of the forecast from the latest orders, shipments or point-of-sale data. Exception detection flags forecasts, orders or inventory positions that look wrong so planners review only those lines. Both are often built on statistics or business rules.
AI that works with language or carries out tasks:
Large language model (LLM) copilots. Chat interfaces built on models like the ones behind ChatGPT that let planners ask questions of their data. A language model is trained to predict the next word in a sequence (its objective) but knows nothing of your specific network, so it is useful for querying data and a poor choice for producing the forecast (more on why).
AI agents. Software that carries out multi-step tasks, such as rebuilding a forecast or raising purchase orders, with varying amounts of human approval. An agent speeds up the work around a forecast and inherits whatever forecast it is given.
AI that reasons about your supply chain:
Reasoning models. A reasoning model is trained to work toward an objective rather than to continue a pattern. In planning, that means a model of how your particular supply chain works, trained on your business and directed at an economic goal such as margin at a target service level. It can evaluate a decision before you make it, including one the history has never seen, such as adding a second supplier or a new distribution center.
When a vendor says its product uses AI, ask which of these it means and which planning decision it changes.
Why does context decide how far AI can take a plan?
Every technology above can only work with what it is given. We call the distance between what a model receives and what it needs to plan well the Context Gap. Closing it takes three kinds of input:
History: what happened. Sales, orders, inventory, pricing, promotions, production and purchasing. It reflects real demand and also the decisions that constrained fulfillment.
Structure: how the business works. The network, lead times, capacity, material compatibility, substitution and product relationships.
Knowledge: what people know is changing. Customer intentions, competitor actions, supplier problems, launch delays and growth plans.
Forecasting methods that predict from history use the first input, and machine learning adds a few drivers such as price. A language model wrapper brings no knowledge of your business. An agent works with whatever forecast it receives. Structure and knowledge usually enter the plan only as manual edits after the forecast is finished.
History on its own carries a problem that better algorithms cannot fix. It records what you were able to do. If a distribution center ran out of stock, the sales history shows low demand. If you capped allocation to a retailer, the history shows the cap. Fit a model to that record and you forecast your constrained self. Buy to that forecast and you constrain the business again, so the next fit is tighter still. Growth gets planned out of the business one cycle at a time.
Cleaning history for stockouts helps a model separate low demand from no stock to sell. It cannot show demand the business never offered to serve, such as the volume a capped retailer would have taken. That demand has to come from structure and from what planners know.
Does machine learning forecast demand better than statistical methods?
Sometimes, and by less than most marketing suggests. Machine learning tends to win when a business has many related products and rich data on drivers such as prices and promotions. On short or sparse sales histories, simple statistical methods often do as well or better.
The best public evidence comes from the M competitions, open forecasting contests run since 1982 by Spyros Makridakis, in which research teams forecast the same real-world data and are scored against simple benchmarks. When researchers later tested machine learning methods on 1,045 short monthly series from the third round, simple statistical methods beat them on every accuracy measure (Makridakis, Spiliotis and Assimakopoulos, PLOS ONE, 2018).
The fifth round in these competitions, M5, is the one that most resembles a planner’s job. Teams forecast Walmart unit sales for 30,490 product and store combinations, with prices and calendar events supplied. Machine learning methods took every top position, and the winner was 22.4% more accurate than the best statistical method (Makridakis, Spiliotis and Assimakopoulos, International Journal of Forecasting, 2022). But, looking closer, two other results matter more to a planner. Only 7.5% of teams beat that statistical method at all. And the advantage shrank as forecasts got more detailed, to about 3% at the product-by-store level where replenishment decisions get made.
Seen through the Context Gap, the result makes sense. Every team worked from history plus prices and a calendar. None had lead times, supply constraints or anything a Walmart planner knew. Better algorithms got more out of history, and the gains thinned out where decisions are made. It also means any accuracy claim is only meaningful once you know the level of detail it was measured at and the kind of demand behind it.
How do spreadsheets, APS and reasoning AI platforms compare?
Spreadsheets, advanced planning systems (APS) and planning platforms built on a reasoning model differ most in how the forecasting model is built, how much of the Context Gap it closes and how many futures a team can evaluate before it commits.
Question | Spreadsheets | Advanced planning system (APS) | Reasoning AI planning platform |
|---|---|---|---|
How is the forecasting model built? | Formulas the planner writes, often moving averages or last year plus growth | A fixed library of model families (exponential smoothing, ARIMA, machine learning ensembles) selected per SKU, the same library for every customer | A model of your supply chain trained on your business and directed at your business objectives |
Which inputs does it use? | History, plus whatever the planner types in | History for the forecast; structure sits in separate supply planning modules and knowledge is applied as manual adjustments | History, structure and knowledge as inputs to the model |
How does planner knowledge get in? | Typed into cells; much of it stays in the planner’s head | Manual adjustments and overrides on top of the statistical forecast | As structured inputs the model uses, with their effect visible in the plan |
How do plans improve? | With planner diligence, up to the limit of the model and history | With planner diligence, up to the limit of the model and history | Automatically through self-learning capabilities |
How many scenarios before a commitment? | Two or three, built by hand | Several, usually configured by an analyst | Hundreds, generated and compared automatically |
How quickly can you ask questions of the data? | Pivot tables and lookups; slow for multi-level questions | Reports and dashboards defined in advance | Questions asked in plain language across the full dataset |
Can you compare a rejected scenario with actuals? | Rarely; old versions get overwritten | Possible with snapshots, seldom done | Kept as part of the record of each plan |
What happens when the network changes? | The workbook gets rebuilt | The model is reconfigured, often as a services project | The model is retrained on the new structure |
Where does it fit best? | Small catalogs, few locations, stable demand | Large organizations that need standard processes and deep ERP integration | Any organization that wants to drive P&L and balance sheet impacts through supply chain optimization |
Why spreadsheets still run so much demand planning
Spreadsheets were adopted for good reasons. They are flexible and fit exactly how one planner thinks, and a skilled planner with a good workbook frequently beats a planning tool, if they had one.
But the deeper problem is that a spreadsheet lets you consider one future at a time. A planner in Excel can model two or three scenarios before Thursday’s meeting and cannot go back afterward to compare the rejected scenario with what actually happened. Every decision gets made on a sample size of one.
The spreadsheet is also where much of the planner’s knowledge lives, and a replacement that cannot absorb it will perform worse than the workbook it replaced. More on this in Why Spreadsheets Aren’t the Answer to Demand Forecasting.
Where advanced planning systems reach their ceiling
An APS brings what spreadsheets lack: shared data, workflow, governance, ERP integration and a forecasting engine that can run thousands of SKUs. Most APS vendors have added machine learning methods in recent years, and for many large organizations the process discipline alone is worth the investment.
The ceiling sits in the model library. Your vendor ships a fixed library of model families, identical for you and every other customer, so the achievable accuracy was set before the project started. Once each SKU has been matched to the best model the library offers, further tuning rarely improves accuracy. This design was state of the art ten years ago, when training a model on each company’s own network was not practical. Planners compensate by adjusting the output by hand, which brings their knowledge back in through the least measurable route. We explore this in Every Supply Chain Needs a Different Forecast.
Is faster planning the same as better planning?
No. Most of what is sold as AI planning today makes the existing decision faster: automate manual tasks, eliminate forgotten tasks and mistakes. That saves real time. It leaves open the question that matters most, whether the plan the business commits to was the best one available.
The value of automation has a cap, set by the size of the planning payroll. But there is a second approach which is now possible with reasoning AI systems: Optimization. Optimization works on a far larger base: the inventory planners position, the freight they expedite, the capacity they book and the sales lost when the plan is wrong. Better decisions reach every asset the planners touch, and the gain does not depend on removing a single planning role. An agent running on a forecast built from history alone also reaches the same ceiling, only sooner.
Two Monday mornings
In the first, an agent rebuilt the forecast overnight, flagged the exceptions and drafted replenishment orders against the new numbers. The planner reviews and approves. The result is the plan the team would have produced by hand on Wednesday, delivered two days early.
In the second, the planning platform evaluated hundreds of versions of next quarter overnight: different promotion depths at the largest retail accounts, a second supplier for a constrained component, a pre-build ahead of a tariff change, launch inventory held at fewer distribution centers. It ranked them against the margin and service targets the business set and surfaced the few worth discussing, including one nobody had proposed. The planner spends the morning choosing among modeled trade-offs, and the team can later compare the rejected options with what actually happened.
Both mornings involve AI. Only the second changes what the business decides.
What does Gartner expect from AI in supply chain planning?
Gartner’s recent research makes three points relevant here:
Spending on automation is running well ahead of results. 83% of senior supply chain leaders surveyed at $500 million-plus organizations had already spent at least $3 million automating planning, yet Gartner predicts only 5% of organizations implementing planning automation will make even 10% of planning decisions autonomously by 2030 (Gartner, September 2026).
AI forecasting will become standard, and data is the main obstacle. Gartner predicts 70% of large organizations will use AI-based forecasting to predict demand by 2030, and lists incomplete data among the main barriers (Gartner, September 2025).
The measure that matters is decision quality. Gartner recommends that leaders measure decision improvement rather than technology deployment, and expects planners to shift from “data administrators” to “plan orchestrators” (Gartner, September 2026).
Read together, they suggest the technology will be widely available, and the difference will come from the inputs it receives and the decisions it improves.
How should planner knowledge get into the plan?
The constraint that decides next quarter is often in a planner’s head and nowhere in the ERP: the retailer about to reset its shelves or the launch slipping a month.
In a 2009 study in the International Journal of Forecasting, Fildes, Goodwin, Lawrence and Nikolopoulos analyzed more than 60,000 forecasts and outcomes from four supply chain companies. In three of the companies, planners’ judgmental adjustments improved accuracy on average. Larger adjustments tended to help. Smaller ones often made accuracy worse, and upward adjustments were wrong in direction more often than downward ones, which the authors read as a bias toward optimism.
Planner judgment carries real information, especially when something significant is changing. Applied as a manual edit to a finished number, it also carries bias nobody measures. A better design treats that knowledge as an input the model can use and the team can audit afterward.
What data do you need for AI demand planning?
A reasonable amount of information, organized around the three inputs of the Context Gap.
History
Order and shipment history at the level of detail you plan, such as SKU by location or SKU by customer, covering your seasonal cycles.
Inventory and stockout records, so the model can tell low demand apart from no stock to sell.
Price and promotion history, including trade promotions and retailer events.
Structure
Locations, lead times, capacities and sourcing rules.
Bills of materials, substitutions and product relationships where they apply.
Knowledge
What planners and commercial teams know about customers, suppliers, launches and growth plans, captured in a structured form going forward rather than reconstructed from old overrides.
Most teams have the history. Fewer have structure in a form a model can use, and very few capture knowledge anywhere outside planners’ heads and spreadsheets. You usually do not need a finished data lake or perfectly clean master data first.
What should you ask about any AI demand planning claim?
Use these six in a vendor demo, an RFP or a review of the platform you already run.
Which kind of AI is this, and which decision does it change? If the answer is only “the forecast arrives faster,” you are buying automation.
Was the model built for our business, or is it a fixed library of model families fitted to our history? Ask what runs underneath, by name.
Which of history, structure and knowledge does the model use directly? Anything that enters only as a manual adjustment is outside the model.
How does a planner’s knowledge get in, and what happens when it contradicts the history? Look for a mechanism you can audit afterward.
How many scenarios can we evaluate before a commitment meeting, and can we compare a rejected one with what actually happened? Without both, the team is still planning one future at a time and cannot learn which judgment calls paid off.
What accuracy do you reach at our planning level of detail, on a holdout of our hardest SKUs, against our current forecast and a simple statistical baseline? An accuracy figure without a stated level cannot be compared with anything, and the M5 results show how hard a simple baseline can be to beat.
Where Omnifold fits
Omnifold builds reasoning AI for supply chain planning. You receive your planning platform, containing your model, developed exclusively for your business: a reasoning model of your supply chain, trained with reinforcement learning and directed at an economic objective. It is a different family from statistical methods such as Holt-Winters, ARIMA and Prophet, from supervised machine learning such as gradient boosting, neural networks and tabular foundation models, and from a language model writing a forecast. Our approach explains how the model is built.
The platform is designed to close the Context Gap. Your model learns from your history and your network’s structure. Planners supply what they know as Enhancements, the platform proposes Enhancements of its own, and the team can see the effect of each one in the plan. Conversations let planners ask questions of the data in plain language. Your data is used only to train a model for you and is not used for any other purpose. See the product for how this works in demand planning.
Ricoh, a $14 billion printer and copier manufacturer, sells through large, infrequent enterprise deals: the sparse demand where machine learning has shown the smallest gains. In a six-week pilot, Omnifold’s unadjusted forecast exceeded 80% accuracy at SKU by customer level. The incumbent’s forecast, after planner adjustments, reached 65% at the coarser SKU level.
Frequently asked questions
Will AI replace demand planners?
No. AI will take over much of the routine work in demand planning, such as data cleansing and exception triage, which frees planners for the decisions that need judgment. Planners who hold commercial and operational knowledge become more valuable, because that knowledge is an input AI cannot get anywhere else. Gartner predicts that by 2030 only 5% of organizations implementing planning automation will make even 10% of planning decisions autonomously (Gartner, September 2026).
What is the difference between AI demand forecasting and AI demand planning?
AI demand forecasting predicts future demand by SKU, location, customer and period. AI demand planning uses that prediction, along with inventory, capacity and commercial knowledge, to decide what to make, buy, stock and promise.
Is machine learning the same as AI in demand planning?
Machine learning forecasting is the technology most often sold as AI demand planning. It learns patterns from sales history and drivers such as price, and it can beat statistical methods on large, data-rich catalogs. It predicts from history and does not reason about your network’s structure or use what your planners know unless that is added by hand.
How accurate is AI demand forecasting?
It depends on the data and on the level of detail it is measured at. In the M5 competition, the winning machine learning method beat the best statistical benchmark by 22.4% overall, and the top methods averaged only about 3% better at the product-by-store level. Ask for accuracy at the level you plan, measured on held-out data from your own business.
Can we use ChatGPT or another large language model for demand forecasting?
Large language models are useful for asking questions of planning data and explaining results. They are a poor choice for producing the forecast, because they are trained to predict text and have no knowledge of your network, lead times or constraints.
Related reading
Every Supply Chain Needs a Different Forecast. Your APS Gives Everyone the Same One.
Sources
Makridakis, S., Spiliotis, E. and Assimakopoulos, V. (2018). Statistical and Machine Learning forecasting methods: Concerns and ways forward. PLOS ONE.
Makridakis, S., Spiliotis, E. and Assimakopoulos, V. (2022). M5 accuracy competition: Results, findings, and conclusions. International Journal of Forecasting.
Fildes, R., Goodwin, P., Lawrence, M. and Nikolopoulos, K. (2009). Effective forecasting and judgmental adjustments. International Journal of Forecasting.
Gartner (September 2026). Gartner Predicts Only 5% of Organizations Will Make At Least 10% of Supply Chain Planning Decisions Autonomously by 2030.
Gartner (September 2025). Gartner Predicts 70% of Large Organizations Will Adopt AI-Based Supply Chain Forecasting to Predict Future Demand by 2030.
