Your Manual to Empower Your Business with AI: How to Build an Agentic ERP
A practical manual for finding valuable AI opportunities, checking complexity and risk, running a controlled pilot, measuring results and turning ERP into a useful second brain for your team.
Short answer: an ERP becomes agentic when it does more than store what happened. It can understand a request in context, retrieve authorised evidence, analyse it, propose the next step and sometimes take a bounded action after clear approval. Think of it as a second brain working with your team: it remembers, connects, calculates, alerts and prepares decisions. It does not own your people’s expertise or accountability, and it should never operate without boundaries or an audit trail.
The strongest project does not start with “Where can we add AI?” It starts with: Where does work lose time, wait for information or repeat the same decision? You then choose one measurable case, test it on real data, compare it with a baseline and expand only when the evidence supports expansion.
- DiscoverFind a real work gapRepetition, waiting, long searches, manual planning or decisions delayed by scattered data.
- AssessUnderstand complexity and riskData quality, exceptions, integration, permissions, error cost and reversibility.
- PilotRun a bounded experimentOne role, one workflow, a short period, a clear sample and human approval for important actions.
- MeasureCompare before and afterTime, quality, cost, business outcome, human intervention and near-miss errors.
What is an agentic ERP?
A traditional ERP records transactions, applies rules and displays reports. An AI agent adds an operating loop:
- Observe: a request, event, exception or changing KPI.
- Understand context: role, customer, product, time, policy and permissions.
- Collect evidence: from authorised ERP data, documents and sources.
- Propose a plan: explain what it found, what it recommends and why.
- Use a tool: run a calculation, optimiser or search, or prepare a draft transaction.
- Request approval: when the action is financial, sensitive or difficult to reverse.
- Record the outcome: who requested it, what evidence was used, what was approved and what happened.
Not every chatbot is an agent. A tool that only answers general questions is an intelligent interface. A system that can read context, use defined tools and prepare or execute a controlled workflow step is closer to an ERP agent.
| Level | What AI does | Human role |
|---|---|---|
| Intelligent search | Finds information and cites its source | Read and decide |
| Assistant | Summarises, calculates and drafts | Review and edit |
| Recommender | Compares options and proposes an action with reasons | Select an option |
| Approval-gated executor | Prepares and executes after explicit confirmation | Approve before write or send |
| Bounded agent | Handles low-risk actions within limits | Monitor exceptions and retain stop authority |
Step 1: find gaps worth empowering
Do not begin with a list of AI features. Walk alongside the work. Follow a shift, an order or a customer journey from beginning to end. At each step ask:
- Where does someone retype information that already exists?
- Where does the team open several screens and files to answer one question?
- Where does a decision wait for someone to assemble data manually?
- Which task repeats dozens or hundreds of times?
- Where does planning depend on one person’s hard-to-transfer experience?
- Where are there too many combinations to evaluate manually?
- Where is a deviation discovered after the correction window has passed?
- Where does a customer receive a number but cannot visualise the outcome?
Scroll through the main opportunity patterns:
- 01Search and retrievalA simple question requires multiple reports, screens and files.
- 02RepetitionCopying, classifying, summarising and entering data follows the same pattern.
- 03PlanningPeople sequence resources, orders or shifts under many changing constraints.
- 04OptimisationA team must choose the best load, cut, layout or route from many possibilities.
- 05Early detectionDelays, shortages, errors or unusual patterns need to surface before they grow.
- 06PersonalisationCustomer choices and context should produce a relevant experience or proposal.
- 07Decision supportEvidence, alternatives and trade-offs must be assembled before human judgement.
- 08Workflow coordinationTasks, alerts and draft records must cross several departments.
The work-gap card
Write one page for each candidate:
- ProblemWhat actually happens?Describe evidence and examples, not impressions.
- VolumeHow often and how many people?Weekly cases and minutes per case.
- ImpactWhat is being lost?Hours, delays, scrap, shipping cost, sales opportunity or decision quality.
- DecisionWhat will AI do?Search, predict, optimise, generate, recommend or execute?
- EvidenceHow will success be known?A primary KPI, a guardrail KPI and a clear baseline.
- OwnerWho owns the result?A business owner, not technology alone.
Step 2: rank opportunities by value and feasibility
Score each candidate from 1 to 5 on two dimensions.
Potential value
- Frequency and volume.
- Current time and cost.
- Customer, revenue or cash-flow impact.
- Quality, safety or compliance impact.
- Speed at which a result can appear.
Feasibility
- Data exists and can be accessed.
- A correct result can be defined and measured.
- Rules and exceptions are understood.
- Required tools and integrations are available.
- Errors can be detected and reversed.
- A business owner and test users are available.
| Value | Feasibility | Decision |
|---|---|---|
| High | High | Strong early-pilot candidate |
| High | Low | Repair the data or process first, then test |
| Low | High | Possible quick improvement, but not a strategic programme |
| Low | Low | Stop or redefine the problem |
Step 3: check complexity before writing an agent
A task can look simple in a demonstration and contain dozens of exceptions in practice. Use this complexity check:
- 01Input varietyIs the input structured tables, or variable text, images, drawings and documents?
- 02Data qualityAre codes, units, dates and record relationships trustworthy?
- 03System countHow many systems and files must the agent read or write?
- 04ExceptionsWhich rare cases completely change the decision?
- 05Decision speedAre there seconds or hours to review and correct?
- 06Error costIs a mistake inconvenient, financial, legal or operationally dangerous?
- 07ReversibilityCan the action be cancelled and the prior state restored?
- 08ExplainabilityCan the agent show its evidence, constraints and calculations?
- 09Environmental changeDo prices, rules, capacity and demand change frequently?
- 10OwnershipWho approves, monitors and stops the agent?
Step 4: give the agent a job description
Write an operating contract before building:
- Mission: the specific outcome it supports.
- Scope: workflows, products and locations included.
- Knowledge sources: authorised tables, documents and versions.
- Tools: what it may read, calculate, create or change.
- Permissions: least privilege for every tool.
- Decision boundary: what it only suggests and what it may execute.
- Approvals: the person or role approving each sensitive action.
- Stop conditions: when it says “I do not know” or escalates.
- Audit: what must be stored for review.
- Metrics: value, quality, safety and cost.
Step 5: prepare data and context
An agent does not become reliable merely because it can query a database. It needs a dependable context layer:
- Business glossary: what active customer, late order, available stock, margin and shift mean.
- Master data: controlled item, customer, supplier, unit and location identities.
- Row and field permissions: the answer reflects what this user is allowed to see.
- Data time: expose the last refresh, especially for inventory and planning.
- Cited sources: show which records and reports support the answer.
- Quality tests: missing values, mixed units, impossible dates and broken relationships.
- Evaluation set: agreed questions and cases used to test every version.
If the data is weak, AI may help detect and classify the problem. It cannot turn an untrusted source into truth.
Step 6: choose the right engine
Not every gap needs a language model. Match the component to the problem:
| Problem type | Usually suitable engine | Example |
|---|---|---|
| Question over structured data | Governed query with a semantic layer | Which orders are late today? |
| Document extraction | Reading and classification with validation | Capture a purchase order or specification |
| Prediction | Statistical or machine-learning model | Forecast demand or process time |
| Constraint optimisation | Mathematical optimiser | Cutting plan or container loading |
| Text or image generation | Generative model with context and controls | Explanation, draft or internal visualisation |
| Multi-step coordination | Agent using defined tools | Collect evidence, prepare a plan, request approval and create a draft |
Step 7: build a pilot that can fail safely
A pilot is not a miniature version of the whole vision. It tests one question:
Can this solution, with this data, help this role improve this KPI without crossing the risk limit?
A practical pilot scope
- One workflow.
- One team or location.
- Historical data or a sandbox first.
- 20–100 cases covering normal and exceptional work, depending on process volume.
- One accountable outcome owner.
- Long enough to capture the real work cycle—often two to six weeks.
- A shadow mode where AI recommends but does not act, followed by comparison with reality.
These are operating suggestions, not universal standards.
Test four sets, not only the average
- Easy, frequent cases.
- Complex, high-value cases.
- Rare or incomplete cases.
- Cases the agent must refuse or escalate.
Testing only perfect cases tests the demo, not the work.
Step 8: agree on measurement before seeing results
Use a scorecard with four views:
- ValueTime, cost and business outcomeMinutes per case, shipping cost, conversion, throughput or scrap.
- QualityCorrectness and consistencyFirst-time-right rate and variance from expert judgement.
- AdoptionUseful use, not curiosityShare of cases where the team used the recommendation and completed the work.
- SafetyErrors, intervention and escalationRejected suggestions, stopped actions, unauthorised data and “I do not know” cases.
A simple value equation
Estimated monthly value = time saved × the appropriate value of that time + avoided waste or cost + attributable incremental profit − operating and review cost.
Do not convert every saved minute into cash when no cost falls and no extra capacity is used. Recovered time may still matter because it speeds response or frees people for higher-value work.
Design a credible comparison
- Record a baseline before the pilot.
- Compare similar cases, not a quiet season with a peak.
- Use a control group when practical.
- Separate the effect of AI from process or incentive changes.
- Monitor for several weeks after intensive support ends.
What do external benchmarks say?
The following are references, not forecasts for your project:
- 5,179 workers14% average productivity gainA field study of a customer-support assistant found more issues resolved per hour, rising to 34% for novice and lower-skilled workers.
- 758 consultants25.1% fasterOn tasks inside the model’s capability frontier, participants completed 12.2% more tasks and delivered over 40% higher human-rated quality.
- Outside the frontier19 percentage points less accurateOn a task outside that frontier, AI users were less likely to reach the correct answer than people without AI.
- Planning case10–12% more accurate forecastsMcKinsey reports one case with 6–8% lower finished-goods inventory and a 3–5% increase in order fill rate; it is a company case, not an industry average.
Step 9: build control before granting action
NIST frames AI risk management as a continuous cycle of governing, mapping, measuring and managing. Microsoft’s agent guidance emphasises least privilege, per-action authorisation, human approval for high-impact actions and tool-level audit logging.
Apply this minimum control set:
- A distinct agent identity, not a shared general account.
- Least privilege for every tool and data source.
- User and agent authorisation checked on every action.
- Human approval for payments, deletion, production changes, external sends and decisions affecting people.
- Step, time and cost limits for each run.
- Logs for inputs, evidence, tools, outputs and approvals.
- Masking or minimising sensitive data that the task does not need.
- Tests for instructions hidden inside documents or messages.
- A kill switch, an on-call owner and a rollback plan.
- Periodic review because models, data and work change.
Step 10: move from assistant to agent in stages
| Stage | Operating mode | Gate to advance |
|---|---|---|
| 0 — Manual | Current process and baseline | Problem and measurement are known |
| 1 — Shadow | AI recommends without affecting work | Stable results on a representative sample |
| 2 — Assistant | Shows evidence and a draft to the user | Clear quality, acceptance and time benefit |
| 3 — Approval-gated action | Creates or changes after explicit confirmation | Permissions, logs, rollback and safety tests |
| 4 — Bounded autonomy | Executes low-risk cases and escalates the rest | Proven limits, continuous monitoring and a stop owner |
Practical work delivered by FLAPP
These examples show different empowerment patterns. Exact results should be measured in each business, so we describe what the solution does and what should be measured rather than assigning an untested number.
1. From a furniture quotation to an authentic 3D view in the customer’s home
For a furniture workflow, the customer selects required units—bedroom, dining room and supporting pieces. Instead of seeing only item names and a price, the solution uses AI to configure an authentic 3D visualisation of the selected units in a home setting.
This addresses an important gap. The customer is not buying codes and dimensions; they are trying to imagine the result in their life. A visual can help the sales team discuss fit and selection earlier and may reduce hesitation or misunderstanding.
- What should be measured?Commercial effectQuote-to-order conversion, decision time and selection revisions.
- What protects the customer?Representation accuracyState that the image is a visualisation and bind every piece to its selected code, dimensions and finish.
- Where does the human remain?Design and approvalVerify measurements, suitability and promises before confirming the order.
2. AI-assisted sawing plans for a marble planning team
In marble operations, a sawing decision balances slab dimensions, defects, vein direction, customer orders, kerf, waste and delivery priority. The number of possible layouts grows quickly. Human expertise remains valuable, but planning can take time and vary by planner.
FLAPP used AI and planning tools to help the planning section propose a sawing plan under the available constraints. The planner did not disappear; they reviewed alternatives and exceptions and approved the plan rather than starting every calculation from zero.
Useful measures include:
- Slab utilisation.
- Waste area or value.
- Planning time.
- Revisions after approval.
- Priority and due-date compliance.
- Cases where the planner rejected the recommendation, with reasons.
3. Optimising pallet placement inside shipping containers
In logistics, FLAPP helped arrange goods pallets inside a container while considering dimensions, orientation and operating constraints, aiming to increase space utilisation and reduce shipping cost per unit.
Filling visible empty space is not enough. A serious plan also considers weight, stability, unloading sequence, stacking limits, clearances and any safety or compatibility restrictions.
| Goal | Value measure | Guardrail measure |
|---|---|---|
| Use space better | Used volume ÷ usable volume | Load stability and stacking limits |
| Reduce cost | Shipping cost per unit or order | No increase in damage or loading time |
| Use fewer containers | Containers for the same plan | Weight, axle and regulatory compliance |
| Plan faster | Minutes to prepare and approve a load plan | Changes required on the loading floor |
4. An in-system AI assistant that answers instead of making people dig
FLAPP enabled businesses to ask questions inside the ERP in natural language and receive answers grounded in authorised company data. Examples include:
- “Which sales orders are late today, and why?”
- “Which materials may fall below reorder level this week?”
- “Which production orders are stopped, and at what stage?”
- “Compare this month’s purchases with last month and explain the largest change.”
- “Which customers exceeded their credit limit and still have open orders?”
- “Summarise shift performance and show the records behind the summary.”
The value is not chat itself. It is reducing the distance between a question and a decision. To prevent a persuasive but wrong answer, the assistant should show the period, business definition, source records and refresh time, and enforce the user’s permissions.
Make AI a second brain, not a black box
A useful second brain performs five functions:
- Documented memory: retrieves the right record, policy and document with its source.
- Attention: surfaces the exception that matters instead of creating more noise.
- Calculation: compares, plans and optimises with the right engine.
- Explanation: shows why it recommends an outcome and what could change it.
- Controlled action: prepares the next step or executes within a defined boundary.
It should not hide uncertainty. A good answer may be: “There is not enough data,” “two sources conflict,” or “this action requires planning approval.”
A 30–60–90-day roadmap
Days 1–30: discover and select
- Choose one workflow with an owner and baseline.
- Observe work and collect an initial 20–50 real examples.
- Score value, feasibility, complexity and risk.
- Define accepted results, refusal cases and escalation.
- Select the starting autonomy level—usually shadow or assistant.
Days 31–60: build and test
- Clean the minimum required data.
- Connect specific sources with specific permissions.
- Build an evaluation set for normal, difficult and refused cases.
- Run in shadow mode, then with a limited user group.
- Measure time, quality, intervention, errors and cost.
Days 61–90: decide and scale carefully
- Compare results with the baseline and control group.
- Fix the largest failure modes, not only the average score.
- Document the operating contract, approvals, audit and stop control.
- Train users to verify and escalate, not merely prompt.
- Decide to stop, redesign, hold the scope or expand to a nearby workflow.
When should you stop the pilot?
Stop or redesign when:
- No business result is visible despite an impressive interface.
- Users must completely redo the AI output.
- Rare errors can cause harm that cannot be contained.
- The source or reason for an answer cannot be reconstructed.
- Model and review cost exceeds value.
- No one owns the process, permissions or incidents.
- Success requires data you do not have the right to use.
- Every case becomes a new exception with no stable boundary.
Common agentic ERP mistakes
- Buying a model, then searching for a problem.
- Starting with a financial or operational action that is difficult to reverse.
- Calling a general chatbot an agent without tools or trusted context.
- Granting broad permissions to avoid proper integration work.
- Reporting average accuracy while hiding worst cases.
- Counting questions instead of time, decisions and outcomes.
- Ignoring human-review and operating cost.
- Using a language model alone for a mathematical optimisation problem.
- Asking people to approve without evidence or enough time.
- Scaling before the data and workflow are stable.
- Failing to teach users when to reject AI output.
- Shipping without a kill switch or rollback plan.
Pre-launch decision checklist
- ProblemCan the gap be stated in one sentence and tied to a KPI?
- DataDo we know the source, owner, quality, timing and permission?
- BoundaryDo we know what the agent does and refuses?
- EvaluationDid we test normal, exceptional, missing-data and adversarial cases?
- ApprovalDo high-impact actions require an accountable person?
- AuditCan we reconstruct what happened and who approved it?
- RollbackCan the action be reversed or operations returned to safety?
- EconomicsDid we include review, integration, model and support costs?
- PeopleCan users verify, escalate and stop the system?
- DecisionIs there a threshold to scale, redesign or stop?
Frequently asked questions
Does agentic ERP mean AI runs the company alone?
No. It means AI can understand context, use tools and assist a sequence of steps within limits. Autonomy depends on risk, and many cases should remain recommendations or approval-gated actions.
What is the best first use case?
A repeated, costly or slow task with accessible data and a measurable correct outcome, where errors are detectable and reversible. Internal search, drafting or optimising a plan before approval is often safer than an autonomous financial action.
Do we need perfect data?
No, but you need enough reliable data for the selected case and an explicit understanding of what is missing. Start narrowly rather than waiting for a universal cleanup, and never hide weak data behind a confident answer.
How should AI return on investment be measured?
Compare time, quality, cost and business outcome before and after, then subtract model, integration, review and support costs. Use a control group where possible, and do not turn every saved minute into cash without a real operating effect.
Is an Arabic AI assistant enough to make an ERP agentic?
Arabic access helps adoption, but it is not enough. The assistant becomes agentic when it uses trusted context, enforces permissions, calls defined tools, requests approval and audits its action.
Sources and benchmark limitations
- NBER: field study of an AI assistant used by 5,179 support workers.
- Harvard Business School: “jagged technological frontier” experiment with 758 consultants.
- McKinsey: autonomous supply-chain planning case and operating results.
- NIST: AI Risk Management Framework and Generative AI Profile.
- Microsoft: shared responsibility model for AI agents.
The numbers come from different tasks, companies and contexts. They do not predict the result of FLAPP or any individual project. Use them to shape a hypothesis and measurement plan, then rely on a controlled pilot using your own data, people and customers.