ERP Implementation

Your Complete Manual to Prepare Master Data for ERP

A detailed manual for cleaning, coding, approving and migrating products, customers, suppliers, units, warehouses and bills of material—with quality metrics and ongoing governance.

FLAPP industrial engineering team ·29 September 2026 ·25 min read
Your Complete Manual to Prepare Master Data for ERP

Short answer: preparing master data is not copying Excel files into an ERP. It means agreeing on one trusted identity for every item, customer, supplier, location and unit; defining ownership, required fields, validity rules and survivorship when sources conflict; then testing that data inside real business processes before go-live.

If incomplete, duplicated or inconsistent data enters the new ERP, it does not improve. Its errors spread faster: wrong inventory, duplicate purchases, conflicting reports, weak production plans and users who stop trusting the system.

  1. DefineKnow what master data isReused business entities: item, customer, supplier, location, unit, account, asset or employee.
  2. CleanCorrect, do not merely moveRemove duplicates, standardise names and units, complete critical fields and link values to evidence.
  3. TestExercise the process, not only the fileBuy, receive, store, produce, sell and return a complete sample using candidate data.
  4. GovernMake quality continuousOwners, approvals, creation and change rules, KPIs and post-launch reviews.

What is master data—and what is not?

Master data describes relatively stable entities reused by many transactions. An item number, unit and category are master data. A sales order is a transaction. Inventory at the cutover moment is an opening balance, not an item master record.

Data typeExamplesTreatment
Master dataItems, customers, suppliers, locations, units, BOMs, work centresClean, code, own and govern continuously
Reference dataCurrencies, countries, tax codes, document types, statusesControlled list with limited change
Transactional dataSales, purchases, production, movements and invoicesDecide what migrates and what remains archived
Opening balancesStock by location and lot, receivables, payables, cash and WIPTime-stamped snapshot with financial and operational reconciliation
Historical dataPrior transactions, reports and documentsSearchable archive or selective migration based on need
The separation matters because each type requires different validation. The item card can be correct while its opening stock is wrong. A customer balance can be correct while the customer exists in three duplicate records.

Why does master data deserve its own workstream?

Every transaction depends on it. One unit-of-measure error can repeat through purchasing, receiving, inventory, production, costing and sales. A wrong supplier address is inconvenient; a wrong bank account can be a financial risk. A wrong case dimension can damage warehousing and shipping plans.

External numbers illustrate the operating burden, but they are not forecasts for your company:

  1. 82%One or more days every weekRespondents in McKinsey’s 2023 MDM survey spent at least one day per week resolving master-data quality issues.
  2. 66%Manual reviewRespondents used manual review to assess, monitor and manage master-data quality.
  3. 62%No defined integration processSurveyed organizations lacked a well-defined process for integrating new and existing sources.
  4. 59%Do not measure data qualityGartner reports that this share of organizations does not measure quality, obscuring both cost and improvement.
Gartner also states that poor data quality costs organizations at least $12.9 million per year on average, based on 2020 research. This is a cross-industry, cross-size figure—not an estimate for a small or midsize company. Its value is showing that the cost extends beyond cleaning time into decisions, inventory, customers and risk.

GS1 US illustrates error amplification with a case-dimension example: a quarter-inch error in recorded case height can severely distort pallet and truck calculations. Treat this as a supply-chain illustration, not a universal performance benchmark.

Step 1: define scope before opening files

Start from the business processes included in the first ERP phase, then derive the required data. If phase one includes purchasing, inventory, production and sales, the likely domains include:

  • 01Items and servicesRaw material, WIP, finished goods, packaging, spare parts and services.
  • 02Units of measureStock, purchase and sales units, conversions, packaging and decimal precision.
  • 03CustomersLegal and trading identities, branches, shipping, billing, credit and tax.
  • 04SuppliersIdentity, categories, payment terms, currency, tax, approved items and bank details.
  • 05Warehouses and locationsPlant, warehouse, zone, bin, receipt, issue and count points.
  • 06Bills of materialComponents, quantities, scrap, substitutes, version and effective dates.
  • 07Production routingsOperations, sequence, work centre, time, capacity and setup.
  • 08Accounts and dimensionsChart of accounts, cost centres, projects, tax and payment methods.
  • 09AssetsType, number, location, purchase date, life, depreciation and maintenance.
  • 10Users and employeesIdentity, role, department, location, manager and access; isolate sensitive HR data.
Do not attempt to improve every dataset in one project. Prioritise by business value and error risk. An item used on go-live day matters more than a customer inactive for seven years.

Step 2: establish clear governance

Assign roles for every data domain, not merely file owners:

RoleResponsibilityExample
Data ownerDefines meaning, policy and acceptable qualityHead of procurement owns supplier data
Data stewardReviews, creates, corrects and follows exceptions dailyAuthorised employee reviews a supplier request
Domain expertDefines business meaning and exceptionsEngineer validates item specification or BOM
Technical custodianImplements validation, integration, access and auditERP or data team
ApproverApproves high-impact changesFinance approves bank or tax changes
ConsumerUses data and reports operational impactPlanning, warehouse, sales and accounting
Everyone must know:

If you cannot name the owner and approver, the domain is not ready—however tidy its spreadsheet looks.

Step 3: create a data dictionary and a decision for every field

Maintain one dictionary containing:

  1. DefinitionWhat does the field mean?Is lead time confirmation-to-receipt, or creation-to-inspection?
  2. SourceWhere is truth?Contract, supplier, physical measurement, finance, engineering or legacy system?
  3. RuleWhat is accepted?Format, range, list, relationship or a condition based on another field.
  4. OwnershipWho resolves conflict?The person holding the file is not always the decision owner.

Step 4: define coding and naming standards

Codes

Keep codes stable, unique and scalable. Avoid embedding too much meaning that can change. A code such as FG-Cairo-Red-2026 becomes misleading when location, colour or classification changes.

Store characteristics in separate attributes and use the code as a stable identity. If readable codes are operationally important, define a short structure with documented limits and generation rules.

Names

A useful name distinguishes the record in search and print:

Type + distinguishing attribute + size/capacity + material/grade + pack when needed.

Marble slab — Carrara — 20 mm — polished is better than white marble, but do not repeat every attribute inside the name.

Define Arabic and English rules, abbreviations, spacing, symbols and word order. Do not let each department invent a convention.

Status and lifecycle

Do not delete a record that has been used. Use states such as draft, under review, active, purchase-blocked, sales-blocked, blocked and expired. Define the transaction effect of each state.

Step 5: prepare every domain in detail

A. Item and product master

For every item, review:

B. Customers

Do not merge customers on name similarity alone; they may be distinct legal entities. Do not keep three records for one entity merely because branches entered its name differently.

C. Suppliers

Bank-detail changes require dual verification, segregation of duties and a complete audit trail. Deduplication never justifies copying financial details without review.

D. Bills of material

For each product and version:

Test multi-level explosion and ensure that no loop makes a product its own direct or indirect component.

E. Routings and work centres

Do not migrate a historical duration before deciding whether it is standard time, actual average or estimate. That definition changes schedules, capacity and cost.

F. Warehouses and locations

Create a hierarchy: company → site → warehouse → zone → bin when needed. For each location define:

G. Finance, tax and dimensions

Finance must sign off this domain and test a journal, invoice, return and reconciliation—not only review the spreadsheet.

Step 6: inventory sources and define survivorship

Create a source register:

SourcePotential strengthRisk
Legacy systemHistory and actual useOld rules, free text and duplicates
Department spreadsheetOperational detailMultiple copies without ownership or change history
Contract or official documentApproved identity and termsMay be stale or difficult to read automatically
Physical measurementActual weight, dimensions and conditionIncomplete sample or uncalibrated tool
Supplier or customerLatest informationRequires validation before approval
Business expertKnows meaning and exceptionsKnowledge may be undocumented or differ across experts
Define precedence per field, not only per record. The contract may win for legal name, physical measurement for dimensions, finance for payment terms and engineering for specification.

Step 7: profile the current data automatically

Measure before changing:

Build a heat map: domain × quality dimension × business impact. Do not clean the easiest field first; start with the field that threatens the process.

Step 8: measure the right quality dimensions

Gartner lists nine common dimensions. GS1 guidance also emphasises completeness, timeliness, accuracy and physical verification. Use what matters to the process:

  • AccuracyDoes the value match reality or the approved source?
  • CompletenessAre required fields and instances present?
  • ConsistencyIs the same value represented consistently across systems?
  • UniquenessDoes each entity have one record inside the defined scope?
  • ValidityDoes the value satisfy its rule, list and format?
  • TimelinessDoes a change arrive before the data is used?
  • IntegrityAre relationships between records complete and correct?
  • PrecisionAre quantity, weight and conversion recorded with suitable rounding?
  • AccessibilityCan an authorised consumer retrieve it when needed?

Useful formulas

Do not compress everything into one score that hides risk. Quality may be 98%, while the missing 2% contains bank accounts or unit conversions.

Step 9: detect duplicates and build the golden record

Begin with deterministic matching: tax ID, GTIN, registration number, email, phone or a trusted external ID. Add fuzzy matching for names and addresses, but do not auto-merge ambiguous results.

For every duplicate cluster:

  1. Decide whether it is one entity or related entities.
    • Select the surviving record or create a new identity.
    • Choose the winning value per field using source, recency and approval.
    • Preserve old-to-new code mapping.
    • Move relationships and transactions under control.
    • Record who merged, why and when.
    • Test that reporting and balances remain intact.

Step 10: clean the defect and prevent its return

Cleaning without prevention creates a repeating project. Address root causes:

DefectCorrectionPrevention
Inconsistent nameTransform and reviewNaming template and controlled attributes
Duplicate customerControlled match and mergeDuplicate check before creation
Wrong unitEvidence, approval and conversionUnit list, required field and range test
Missing fieldComplete from sourceConditional requirement and request owner
Stale valueUpdate with effective datePeriodic review and expiry alert
Untraceable changeRebuild from evidenceApproval, audit trail and permissions

Step 11: design the load file as a controlled product

Do not use an ad-hoc file that changes every day. Create a versioned template with:

Protect reference columns and formulas, use controlled lists and never circulate bank or personal data in an unprotected file.

Step 12: run migration rehearsals

Do not make the first complete load the production load:

  • Cycle 0Small sampleTest template, coding, fields and basic rules.
  • Cycle 1Early full scopeReveal real duplicate, missing-data and relationship volume.
  • Cycle 2Corrected scopeTest performance, reports, processes and printed documents.
  • Cycle 3Cutover simulationUse the timing, roles and reconciliation expected on launch night.
  • Final cycleFrozen and approved dataLoad, reconcile and sign off before transactions open.
For every cycle, retain source, loaded, rejected, merged and deferred counts. Explain every difference. “The tool showed no error” is not reconciliation.

Step 13: test data through end-to-end scenarios

Select high-risk and high-frequency samples, then execute:

Observe name, unit, quantity, account, tax, cost, report and printed document. The defect may not appear on the item card; it may appear at the end of the chain.

Step 14: separate master data from the cutover snapshot

After master records are approved, prepare the cutover snapshot.

Opening inventory

Open business data

Define a freeze time, transaction handling during cutover and approval for any post-load difference.

Step 15: define acceptance gates

There is no universal quality percentage for every company. Set thresholds by field impact. An example requiring domain-owner approval:

  1. 100%Critical fieldsUnique identity, base unit, tax, control account, approved bank account and active BOM component.
  2. ≥99.5%Active-record completenessSuggested target for mandatory fields before launch, with approved exceptions.
  3. 0Unresolved confirmed duplicatesEvery confirmed cluster has a documented merge, separation or deferral decision.
  4. 100%Financial balance reconciliationLoad-to-source and general-ledger difference is zero or formally explained and approved.
These are suggested operating targets, not Gartner or GS1 industry benchmarks. A safety or pharmaceutical process may demand stricter gates; a low-risk domain may accept more documented exceptions.

Data-owner sign-off

Step 16: govern after go-live

Data quality decays when products, people or systems change without process change. GS1 recommends regular master-file, governance and training reviews and physical product sampling because item accuracy can deteriorate over time.

A useful monthly dashboard includes:

  1. First-time-right creationSource qualityRequests approved without rework ÷ total creation requests.
  2. Request cycle timeService speedFrom request to approval and publication.
  3. Defects per 1,000 recordsQuality trendBy domain, rule and owner—not only a total number.
  4. New duplicatesPrevention effectivenessClusters created after launch and their source.
  5. Expiry and reviewData currencyRecords requiring review or documents approaching expiry.
  6. Business impactValueBlocked transactions, corrections or variances caused by master data.
Review measures with business owners. Correct the rule, training or source—not only the individual record.

An eight-week working plan

Week 1: scope and ownership

Week 2: dictionary and standards

Week 3: profiling and baseline

Weeks 4 and 5: cleansing and approval

Week 6: trial migration and testing

Week 7: cutover simulation

Week 8: final load and handover

Eight weeks is an organising example. A narrow domain may take less; multiple plants or deep BOM structures may take longer. Never remove testing simply to fit the date.

Common master-data mistakes

  1. Migrating every record “just in case.”
    • Leaving data decisions to technology or the vendor alone.
    • Cleaning names while ignoring units, accounts and relationships.
    • Merging records by name only.
    • Encoding too many changeable meanings inside identifiers.
    • Mixing master data, balances and history in one file.
    • Measuring completeness while ignoring real-world accuracy.
    • Assuming a populated field is correct.
    • Running only one migration before go-live.
    • Accepting unexplained count or balance differences.
    • Giving everyone master-data edit access after launch.
    • Ending the cleanup without an ongoing governance process.

How FLAPP experts can help

A FLAPP expert should not take your files and decide their meaning away from your team. The role is to structure decisions, expose defects and translate business rules into templates, ERP controls and tests, while approval remains with your internal data owners.

FLAPP can support:

  • 01Data-discovery workshopsConnect ERP phases to domains, fields, risks and business ownership.
  • 02Dictionary and standardsDefine fields, codes, names, units and reference values.
  • 03Automated profilingMeasure missing data, duplicates, outliers and broken relationships before editing.
  • 04Transformation mappingMap legacy fields and codes to FLAPP with documented rules.
  • 05Cleansing supportRank duplicate clusters and exceptions and route them to the right decision owner.
  • 06Controlled load templatesLists, validation, error messages and review and approval states.
  • 07Migration cyclesRepeated trial loads with rejection, variance, count and reconciliation reports.
  • 08Process testingExercise purchasing, inventory, production, sales and finance on real candidate data.
  • 09Cutover planningFreeze, load, balances, approvals, rollback and temporary transaction handling.
  • 10Post-launch governanceCreate/change access, approvals, KPIs, monitoring and steward training.

What FLAPP needs from you

FLAPP can accelerate profiling, transformation and migration and provide implementation experience. It cannot independently guess which legal identity, unit conversion or BOM represents your factory. The strongest result is a partnership: FLAPP experts structure and execute; your business owners define and approve.

Frequently asked questions

When should master-data preparation begin?

At the start of ERP design, not a few weeks before launch. Rules and ownership take time, and migration rehearsals need early data.

Should every historical transaction move to the new ERP?

Not always. Move what operations, law or analysis requires and consider a searchable archive for the rest. Moving everything adds cost and can transfer errors without value.

Does business or IT own master data?

Business owns meaning, quality and decisions. Technology manages structure, rules, integration and protection. Success requires both.

Can AI clean the data?

It can suggest matches, classifications, standardised names and anomalies. It should not merge sensitive records or decide business truth without evidence and approval. Use AI to accelerate review, not hide it.

What is the most important pre-launch metric?

There is no single metric. Focus on process-critical fields, scenario success, explained count and balance differences and clear ownership of every exception.

Sources and benchmark limitations

External figures cover different organizations, sizes and industries. Do not use a survey percentage or average cost as your project ROI. Measure your own correction time, blocked transactions, inventory, variances, returns and manual work, then establish a baseline and risk-appropriate target.