意思決定インテリジェンス
AI のためのデータ準備:実運用が要求する条件
必要なのは完璧なデータではなく、対象業務に関する所在、権限、鮮度、品質が説明できる状態です。データ整備は AI 導入の前提工程であり、並行して進める設計が現実的です。
Decision intelligenceWritten by NirjiX AI AdvisoryPublished January 2026Last reviewed February 20269 min read
Direct answer
結論
必要なのは完璧なデータではなく、対象業務に関する所在、権限、鮮度、品質が説明できる状態です。データ整備は AI 導入の前提工程であり、並行して進める設計が現実的です。
以下の詳細分析は英語原文のまま掲載しています。 英語版のフルガイドを読む
Estate maturity is the wrong test
Ask whether the enterprise data estate is ready for AI and the answer is always no, in every organization, permanently. Estates contain decades of accumulated systems, and remediating all of them before starting is an argument for never starting.
The useful question is narrower: for this specific use case, is the data we need reachable, permitted, complete enough and understood? That question is answerable within weeks and frequently produces a surprising result — the blocker is authorization or lineage documentation, not quality.
This reframing matters commercially. Estate-level programs consume years of budget before producing a business outcome, and they are usually the first thing cut when priorities shift. Use-case-scoped readiness work delivers value while leaving durable improvements behind.
The four conditions, and how to test each one
Test them in this order. Accessibility and authorization fail more often than quality, and they fail earlier.
| Condition | The practical test | Remedy when it fails |
|---|---|---|
| Accessibility | Can the delivery team obtain a working sample within two weeks through a supported route? | An access pattern and environment, not a new platform. |
| Authorization | Is there a named owner who can approve this use, and does the purpose fall within existing consent and contracts? | Data ownership assignment and a documented purpose review. |
| Sufficiency | Does coverage span the population and period the decision needs, with known gaps quantified? | Backfill, supplementary source, or narrowing the use case scope honestly. |
| Interpretability | Can someone explain what each field means, how it is produced and where it breaks? | Lineage and definition documentation from the owning function, not from IT alone. |
| Freshness | Does the data arrive fast enough for the decision cadence the use case requires? | Pipeline change, or a decision to operate on a slower cadence deliberately. |
Signals that data will block delivery later
- No single named owner for the primary dataset — the most reliable predictor of a stalled use case.
- Definitions that differ between the reporting layer and the operational system that generates the record.
- Historical data that changed meaning after a system migration, with no documented cut-over.
- Quality that is acceptable for aggregate reporting but not for record-level decisions.
- Personal or contractual data whose permitted purposes were never assessed for analytical or model use.
- Manual correction steps applied downstream that the source system never receives back.
These surface late and expensively. Look for them during assessment rather than during deployment.
How to run a use-case data readiness review
Two to three weeks per use case is normally sufficient.
- 01
Write the decision first
State the decision the model will inform, its cadence, and the consequence of being wrong. Readiness standards for a pricing recommendation and for a triage suggestion are not the same, and treating them identically over-engineers one and under-protects the other.
- 02
Trace the data to its source
Follow each required field back to the system that creates it, not the warehouse that stores it. Most interpretation errors originate in the gap between those two.
- 03
Assess a real sample
Measure coverage, completeness and drift on actual records for the period the use case needs. An inventory of tables is not evidence about quality.
- 04
Resolve permission explicitly
Confirm the purpose is permitted under the contracts, notices and consent that apply, and record the conclusion. Discovering this after build is the most expensive sequencing mistake available.
- 05
Record the residual gaps
Publish what remains imperfect and how the design compensates. A use case shipped with documented limitations is defensible; one shipped with undocumented ones is not.
NirjiX view
The NirjiX view
We treat unresolved data ownership as a gate rather than a score. A use case whose primary dataset has no accountable owner is not a low-readiness candidate — it is not yet a candidate, and the work in front of it is governance, not engineering.
We also resist the sequencing that says all data must be fixed before AI begins. In practice the fastest route to a governed estate is a small number of high-value use cases that force ownership, definitions and access patterns to be resolved for the data that actually matters.
Frequently asked executive questions
- Do we need a data lake or lakehouse before starting with AI?
- No. Consolidated storage helps at scale, but it is neither necessary nor sufficient for a first use case. What matters is whether the specific data a use case requires is accessible, authorized, sufficient and understood. Many organizations with mature platforms still fail on authorization, and many without one deliver successfully.
- How good does data quality need to be?
- Good enough for the decision it supports, with the gaps quantified. A recommendation that a human reviews tolerates more noise than an automated action with financial consequence. The failure is not imperfect data — it is imperfect data whose limitations are undocumented and therefore invisible to the people relying on the output.
- Should unstructured data be included in the readiness assessment?
- Yes, and with the same four conditions. Documents, tickets and transcripts frequently carry more of the value than structured tables, but they are also where authorization and retention questions are least likely to have been answered. Assess them explicitly rather than assuming they are unconstrained.
- Who owns data readiness — IT or the business?
- Accountability sits with the business function that produces and uses the data; IT owns the access route and the platform. When readiness is delegated wholly to IT, definitions and permitted-purpose questions go unanswered because the answers were never IT's to give.
- What would change the assessment?
- A change in the decision the model supports. Narrowing a use case from an automated action to a human-reviewed recommendation can move it from blocked to deliverable without any change to the data itself, because the readiness bar is set by consequence.
Continue
Related intelligence
- DecisionAI readinessThe wider organizational test that data readiness sits inside.
- DecisionAI governance and riskRisk tiers and approval paths that determine how much evidence a dataset needs.
- DecisionAI use case prioritizationWhy data readiness is a gate in the scoring rather than one factor among many.
- PortalRun the AI readiness assessmentScore data foundations alongside the other dimensions that decide delivery.
このテーマの中心ガイド
自社がAI導入の準備が整っているかをどう判断するか。
本ページは意思決定の一側面を扱います。AI準備度アセスメントに関するNirjiXの総合ガイドで全体像をご確認ください。
AI準備度アセスメント →関連情報
関連ガイド
同じ領域の関連ガイド
- PoC から本番へ:AI 施策が止まる五つの関門
停滞は予測可能な五つの関門で起きます。実装されなくても損をする事業責任者が不在であること、PoC のデータが手作業で用意され統制下では本番頻度で取得できないこと、合格基準が未定義で承認が閉じないこと、周辺業務プロセスが再設計されていないこと、そして運用の受け皿が決まっていないことです。
もう一方の領域における同じ意思決定
- GCC の拡張:立ち上げ後の進化の道筋
拡張は人員増ではなく、担う業務の複雑度を上げることで進みます。定型処理から分析、設計、意思決定支援へと責任範囲を段階的に移すことが価値の源泉になります。
Transparency
Sources and methodology
This page reflects NirjiX advisory practice rather than a survey or a vendor benchmark. The structure of the assessment — the dimensions, the maturity language and the sequencing logic — is the same framework used inside the NirjiX AI readiness assessment and the AI plan builder.
Where we describe patterns ("most organizations discover…"), we are describing what we observe across client engagements, not a measured statistic. We deliberately avoid quoting market numbers we cannot verify, because an AI investment case built on borrowed statistics collapses the first time a CFO tests it.
Any figure that ends up in your own plan should come from your own data: your cost base, your cycle times, your error rates, your volumes. The assessment and plan builder are designed to force that discipline.
Assess readiness against a real use case
The AI assessment scores data foundations together with ownership, governance and execution capacity, so the blocker is identified before the build is funded.
Outputs are preliminary and intended for advisor validation before funding decisions.