AI & ERP

Your Manual to Empower Your Business with AI: How to Build an Agentic ERP

A practical manual for finding valuable AI opportunities, checking complexity and risk, running a controlled pilot, measuring results and turning ERP into a useful second brain for your team.

FLAPP industrial engineering team ·28 September 2026 ·24 min read
Your Manual to Empower Your Business with AI: How to Build an Agentic ERP

Short answer: an ERP becomes agentic when it does more than store what happened. It can understand a request in context, retrieve authorised evidence, analyse it, propose the next step and sometimes take a bounded action after clear approval. Think of it as a second brain working with your team: it remembers, connects, calculates, alerts and prepares decisions. It does not own your people’s expertise or accountability, and it should never operate without boundaries or an audit trail.

The strongest project does not start with “Where can we add AI?” It starts with: Where does work lose time, wait for information or repeat the same decision? You then choose one measurable case, test it on real data, compare it with a baseline and expand only when the evidence supports expansion.

  1. DiscoverFind a real work gapRepetition, waiting, long searches, manual planning or decisions delayed by scattered data.
  2. AssessUnderstand complexity and riskData quality, exceptions, integration, permissions, error cost and reversibility.
  3. PilotRun a bounded experimentOne role, one workflow, a short period, a clear sample and human approval for important actions.
  4. MeasureCompare before and afterTime, quality, cost, business outcome, human intervention and near-miss errors.

What is an agentic ERP?

A traditional ERP records transactions, applies rules and displays reports. An AI agent adds an operating loop:

  1. Observe: a request, event, exception or changing KPI.
    • Understand context: role, customer, product, time, policy and permissions.
    • Collect evidence: from authorised ERP data, documents and sources.
    • Propose a plan: explain what it found, what it recommends and why.
    • Use a tool: run a calculation, optimiser or search, or prepare a draft transaction.
    • Request approval: when the action is financial, sensitive or difficult to reverse.
    • Record the outcome: who requested it, what evidence was used, what was approved and what happened.

Not every chatbot is an agent. A tool that only answers general questions is an intelligent interface. A system that can read context, use defined tools and prepare or execute a controlled workflow step is closer to an ERP agent.

LevelWhat AI doesHuman role
Intelligent searchFinds information and cites its sourceRead and decide
AssistantSummarises, calculates and draftsReview and edit
RecommenderCompares options and proposes an action with reasonsSelect an option
Approval-gated executorPrepares and executes after explicit confirmationApprove before write or send
Bounded agentHandles low-risk actions within limitsMonitor exceptions and retain stop authority
Start with the lowest level that delivers value. Autonomy is not a prize; it is added responsibility and risk.

Step 1: find gaps worth empowering

Do not begin with a list of AI features. Walk alongside the work. Follow a shift, an order or a customer journey from beginning to end. At each step ask:

Scroll through the main opportunity patterns:

  • 01Search and retrievalA simple question requires multiple reports, screens and files.
  • 02RepetitionCopying, classifying, summarising and entering data follows the same pattern.
  • 03PlanningPeople sequence resources, orders or shifts under many changing constraints.
  • 04OptimisationA team must choose the best load, cut, layout or route from many possibilities.
  • 05Early detectionDelays, shortages, errors or unusual patterns need to surface before they grow.
  • 06PersonalisationCustomer choices and context should produce a relevant experience or proposal.
  • 07Decision supportEvidence, alternatives and trade-offs must be assembled before human judgement.
  • 08Workflow coordinationTasks, alerts and draft records must cross several departments.

The work-gap card

Write one page for each candidate:

  1. ProblemWhat actually happens?Describe evidence and examples, not impressions.
  2. VolumeHow often and how many people?Weekly cases and minutes per case.
  3. ImpactWhat is being lost?Hours, delays, scrap, shipping cost, sales opportunity or decision quality.
  4. DecisionWhat will AI do?Search, predict, optimise, generate, recommend or execute?
  5. EvidenceHow will success be known?A primary KPI, a guardrail KPI and a clear baseline.
  6. OwnerWho owns the result?A business owner, not technology alone.

Step 2: rank opportunities by value and feasibility

Score each candidate from 1 to 5 on two dimensions.

Potential value

Feasibility

ValueFeasibilityDecision
HighHighStrong early-pilot candidate
HighLowRepair the data or process first, then test
LowHighPossible quick improvement, but not a strategic programme
LowLowStop or redefine the problem
Do not select the most impressive demonstration. Select the case with visible pain, usable data, a measurable outcome and a containable failure mode.

Step 3: check complexity before writing an agent

A task can look simple in a demonstration and contain dozens of exceptions in practice. Use this complexity check:

  • 01Input varietyIs the input structured tables, or variable text, images, drawings and documents?
  • 02Data qualityAre codes, units, dates and record relationships trustworthy?
  • 03System countHow many systems and files must the agent read or write?
  • 04ExceptionsWhich rare cases completely change the decision?
  • 05Decision speedAre there seconds or hours to review and correct?
  • 06Error costIs a mistake inconvenient, financial, legal or operationally dangerous?
  • 07ReversibilityCan the action be cancelled and the prior state restored?
  • 08ExplainabilityCan the agent show its evidence, constraints and calculations?
  • 09Environmental changeDo prices, rules, capacity and demand change frequently?
  • 10OwnershipWho approves, monitors and stops the agent?
As error cost rises and reversibility falls, reduce autonomy and increase testing and human approval.

Step 4: give the agent a job description

Write an operating contract before building:

Step 5: prepare data and context

An agent does not become reliable merely because it can query a database. It needs a dependable context layer:

  1. Business glossary: what active customer, late order, available stock, margin and shift mean.
    • Master data: controlled item, customer, supplier, unit and location identities.
    • Row and field permissions: the answer reflects what this user is allowed to see.
    • Data time: expose the last refresh, especially for inventory and planning.
    • Cited sources: show which records and reports support the answer.
    • Quality tests: missing values, mixed units, impossible dates and broken relationships.
    • Evaluation set: agreed questions and cases used to test every version.

If the data is weak, AI may help detect and classify the problem. It cannot turn an untrusted source into truth.

Step 6: choose the right engine

Not every gap needs a language model. Match the component to the problem:

Problem typeUsually suitable engineExample
Question over structured dataGoverned query with a semantic layerWhich orders are late today?
Document extractionReading and classification with validationCapture a purchase order or specification
PredictionStatistical or machine-learning modelForecast demand or process time
Constraint optimisationMathematical optimiserCutting plan or container loading
Text or image generationGenerative model with context and controlsExplanation, draft or internal visualisation
Multi-step coordinationAgent using defined toolsCollect evidence, prepare a plan, request approval and create a draft
A useful agent can orchestrate several tools. Do not ask a conversational model to guess an optimisation result that a mathematical engine can calculate.

Step 7: build a pilot that can fail safely

A pilot is not a miniature version of the whole vision. It tests one question:

Can this solution, with this data, help this role improve this KPI without crossing the risk limit?

A practical pilot scope

These are operating suggestions, not universal standards.

Test four sets, not only the average

  1. Easy, frequent cases.
    • Complex, high-value cases.
    • Rare or incomplete cases.
    • Cases the agent must refuse or escalate.

Testing only perfect cases tests the demo, not the work.

Step 8: agree on measurement before seeing results

Use a scorecard with four views:

  1. ValueTime, cost and business outcomeMinutes per case, shipping cost, conversion, throughput or scrap.
  2. QualityCorrectness and consistencyFirst-time-right rate and variance from expert judgement.
  3. AdoptionUseful use, not curiosityShare of cases where the team used the recommendation and completed the work.
  4. SafetyErrors, intervention and escalationRejected suggestions, stopped actions, unauthorised data and “I do not know” cases.

A simple value equation

Estimated monthly value = time saved × the appropriate value of that time + avoided waste or cost + attributable incremental profit − operating and review cost.

Do not convert every saved minute into cash when no cost falls and no extra capacity is used. Recovered time may still matter because it speeds response or frees people for higher-value work.

Design a credible comparison

What do external benchmarks say?

The following are references, not forecasts for your project:

  1. 5,179 workers14% average productivity gainA field study of a customer-support assistant found more issues resolved per hour, rising to 34% for novice and lower-skilled workers.
  2. 758 consultants25.1% fasterOn tasks inside the model’s capability frontier, participants completed 12.2% more tasks and delivered over 40% higher human-rated quality.
  3. Outside the frontier19 percentage points less accurateOn a task outside that frontier, AI users were less likely to reach the correct answer than people without AI.
  4. Planning case10–12% more accurate forecastsMcKinsey reports one case with 6–8% lower finished-goods inventory and a 3–5% increase in order fill rate; it is a company case, not an industry average.
The lesson is not to multiply your business case by these percentages. The lesson is that task fit determines value: AI can raise performance inside its capabilities and create confident error outside them.

Step 9: build control before granting action

NIST frames AI risk management as a continuous cycle of governing, mapping, measuring and managing. Microsoft’s agent guidance emphasises least privilege, per-action authorisation, human approval for high-impact actions and tool-level audit logging.

Apply this minimum control set:

Step 10: move from assistant to agent in stages

StageOperating modeGate to advance
0 — ManualCurrent process and baselineProblem and measurement are known
1 — ShadowAI recommends without affecting workStable results on a representative sample
2 — AssistantShows evidence and a draft to the userClear quality, acceptance and time benefit
3 — Approval-gated actionCreates or changes after explicit confirmationPermissions, logs, rollback and safety tests
4 — Bounded autonomyExecutes low-risk cases and escalates the restProven limits, continuous monitoring and a stop owner
Stage two may be the correct final product. Not every process needs stage four.

Practical work delivered by FLAPP

These examples show different empowerment patterns. Exact results should be measured in each business, so we describe what the solution does and what should be measured rather than assigning an untested number.

1. From a furniture quotation to an authentic 3D view in the customer’s home

For a furniture workflow, the customer selects required units—bedroom, dining room and supporting pieces. Instead of seeing only item names and a price, the solution uses AI to configure an authentic 3D visualisation of the selected units in a home setting.

This addresses an important gap. The customer is not buying codes and dimensions; they are trying to imagine the result in their life. A visual can help the sales team discuss fit and selection earlier and may reduce hesitation or misunderstanding.

  1. What should be measured?Commercial effectQuote-to-order conversion, decision time and selection revisions.
  2. What protects the customer?Representation accuracyState that the image is a visualisation and bind every piece to its selected code, dimensions and finish.
  3. Where does the human remain?Design and approvalVerify measurements, suitability and promises before confirming the order.
Do not say “conversion increased” before comparing against a baseline and a suitable group. Say the solution was designed to close the imagination gap, then measure the effect.

2. AI-assisted sawing plans for a marble planning team

In marble operations, a sawing decision balances slab dimensions, defects, vein direction, customer orders, kerf, waste and delivery priority. The number of possible layouts grows quickly. Human expertise remains valuable, but planning can take time and vary by planner.

FLAPP used AI and planning tools to help the planning section propose a sawing plan under the available constraints. The planner did not disappear; they reviewed alternatives and exceptions and approved the plan rather than starting every calculation from zero.

Useful measures include:

3. Optimising pallet placement inside shipping containers

In logistics, FLAPP helped arrange goods pallets inside a container while considering dimensions, orientation and operating constraints, aiming to increase space utilisation and reduce shipping cost per unit.

Filling visible empty space is not enough. A serious plan also considers weight, stability, unloading sequence, stacking limits, clearances and any safety or compatibility restrictions.

GoalValue measureGuardrail measure
Use space betterUsed volume ÷ usable volumeLoad stability and stacking limits
Reduce costShipping cost per unit or orderNo increase in damage or loading time
Use fewer containersContainers for the same planWeight, axle and regulatory compliance
Plan fasterMinutes to prepare and approve a load planChanges required on the loading floor

4. An in-system AI assistant that answers instead of making people dig

FLAPP enabled businesses to ask questions inside the ERP in natural language and receive answers grounded in authorised company data. Examples include:

The value is not chat itself. It is reducing the distance between a question and a decision. To prevent a persuasive but wrong answer, the assistant should show the period, business definition, source records and refresh time, and enforce the user’s permissions.

Make AI a second brain, not a black box

A useful second brain performs five functions:

  1. Documented memory: retrieves the right record, policy and document with its source.
    • Attention: surfaces the exception that matters instead of creating more noise.
    • Calculation: compares, plans and optimises with the right engine.
    • Explanation: shows why it recommends an outcome and what could change it.
    • Controlled action: prepares the next step or executes within a defined boundary.

It should not hide uncertainty. A good answer may be: “There is not enough data,” “two sources conflict,” or “this action requires planning approval.”

A 30–60–90-day roadmap

Days 1–30: discover and select

Days 31–60: build and test

Days 61–90: decide and scale carefully

When should you stop the pilot?

Stop or redesign when:

Common agentic ERP mistakes

  1. Buying a model, then searching for a problem.
    • Starting with a financial or operational action that is difficult to reverse.
    • Calling a general chatbot an agent without tools or trusted context.
    • Granting broad permissions to avoid proper integration work.
    • Reporting average accuracy while hiding worst cases.
    • Counting questions instead of time, decisions and outcomes.
    • Ignoring human-review and operating cost.
    • Using a language model alone for a mathematical optimisation problem.
    • Asking people to approve without evidence or enough time.
    • Scaling before the data and workflow are stable.
    • Failing to teach users when to reject AI output.
    • Shipping without a kill switch or rollback plan.

Pre-launch decision checklist

  • ProblemCan the gap be stated in one sentence and tied to a KPI?
  • DataDo we know the source, owner, quality, timing and permission?
  • BoundaryDo we know what the agent does and refuses?
  • EvaluationDid we test normal, exceptional, missing-data and adversarial cases?
  • ApprovalDo high-impact actions require an accountable person?
  • AuditCan we reconstruct what happened and who approved it?
  • RollbackCan the action be reversed or operations returned to safety?
  • EconomicsDid we include review, integration, model and support costs?
  • PeopleCan users verify, escalate and stop the system?
  • DecisionIs there a threshold to scale, redesign or stop?

Frequently asked questions

Does agentic ERP mean AI runs the company alone?

No. It means AI can understand context, use tools and assist a sequence of steps within limits. Autonomy depends on risk, and many cases should remain recommendations or approval-gated actions.

What is the best first use case?

A repeated, costly or slow task with accessible data and a measurable correct outcome, where errors are detectable and reversible. Internal search, drafting or optimising a plan before approval is often safer than an autonomous financial action.

Do we need perfect data?

No, but you need enough reliable data for the selected case and an explicit understanding of what is missing. Start narrowly rather than waiting for a universal cleanup, and never hide weak data behind a confident answer.

How should AI return on investment be measured?

Compare time, quality, cost and business outcome before and after, then subtract model, integration, review and support costs. Use a control group where possible, and do not turn every saved minute into cash without a real operating effect.

Is an Arabic AI assistant enough to make an ERP agentic?

Arabic access helps adoption, but it is not enough. The assistant becomes agentic when it uses trusted context, enforces permissions, calls defined tools, requests approval and audits its action.

Sources and benchmark limitations

The numbers come from different tasks, companies and contexts. They do not predict the result of FLAPP or any individual project. Use them to shape a hypothesis and measurement plan, then rely on a controlled pilot using your own data, people and customers.