- Key takeaways
- What data strategy should enterprises use to support AI workloads?
- What makes enterprise data AI-ready?
- How can enterprises reduce unnecessary data movement for AI workloads?
- How can companies improve data accessibility for AI while maintaining governance and control?
- How do data pipelines and orchestration help companies scale AI?
Key takeaways
- A strong enterprise AI data strategy gives workloads access to trusted data without requiring every dataset to be copied into one location.
- AI-ready data must be accurate, current, secure, discoverable, and usable by authorized applications.
- Processing data closer to where it already resides can reduce unnecessary transfers and help organizations maintain control.
- Governed data pipelines and orchestration make data preparation repeatable as AI projects move from pilots into production. Dell AI Data Platform with NVIDIA provides a governed data foundation within the Dell AI Factory with NVIDIA, combining Dell data infrastructure and orchestration with NVIDIA acceleration.
Enterprises should build an enterprise AI data strategy that makes trusted, governed data available where AI workloads need it while reducing unnecessary data movement. AI-ready data is only part of the equation. Enterprises also need controlled access and a repeatable way to prepare and deliver information as AI projects scale.
Enterprise data is already spread across more locations. Dell Technologies, citing Gartner research, reports that 75% of enterprise-managed data is created and processed outside traditional data centers. As more information is generated at the edge and across distributed environments, enterprises need to decide whether the data should move or the processing should move closer to it.
What data strategy should enterprises use to support AI workloads?
Enterprises should design their AI data strategy around the requirements of each workload. For every AI use case, teams should first identify the data the workload needs and the rules governing access. They can then decide how current that data must be and where processing should occur.
Gartner recommends aligning data to specific AI use cases and identifying the data sources and governance requirements each one needs.
The Dell AI Data Platform with NVIDIA supports data discovery, preparation, search, and governed orchestration no matter where data sits. Its data and orchestration engines can help teams find, prepare, enrich, and connect structured, semi-structured, and unstructured information into repeatable governed pipelines for AI projects. The platform is designed to keep data where it is governed and move only what the workload requires.
NVIDIA’s technology accelerates that data foundation through GPU-accelerated computing, networking, and AI software. In the Dell AI Data Platform with NVIDIA, technologies such as NVIDIA RAPIDS, cuVS, NVIDIA BlueField data processing units (DPUs), and NVIDIA AI Enterprise can help accelerate data preparation, indexing, retrieval, and AI application deployment while maintaining enterprise controls.
What makes enterprise data AI-ready?
AI-ready data is information an authorized workload can find, understand, and use in the form it needs. That means the data has to be trustworthy and current, but it also has to be easy for the right application to discover and access. A dataset may be technically available and still be difficult to use if an application cannot determine whether it is reliable or appropriate for the task.
AI-ready data should be:
- Accurate and current: Reliable enough for the intended workload.
- Discoverable: Easy for authorized applications to find and understand.
- Secure: Accessible only to approved users and workloads.
- Traceable: Supported by metadata and lineage.
- Fit for purpose: Available in a format the use case can actually use.
Enterprise data feeding AI applications is often unstructured, fragmented, and rapidly changing, making preparation harder. The Dell AI Data Platform with NVIDIA addresses that challenge by combining Dell’s data engines and storage engines with NVIDIA-accelerated infrastructure and AI software. Together, these capabilities help organizations discover, prepare, and activate enterprise data for AI applications in real time.
Many organizations still have work to do before their data can support new AI use cases. Gartner reports that 90% of chief data and analytics officers (CDAOs) say their current data architecture needs an overhaul to support new AI use cases. The Dell AI Data Platform addresses these requirements with capabilities for discovery, preparation, search, and governed pipelines. Teams can use these capabilities to turn raw enterprise information into datasets suitable for models and AI applications.
How can enterprises reduce unnecessary data movement for AI workloads?
Enterprises can reduce unnecessary data movement by processing information closer to where it already resides when the workload allows it. Some data will still need to be copied, transformed, or transferred, especially when applications combine multiple sources.
Before moving data, teams should first determine whether the workload actually needs a local copy. They should also weigh how current the information must be against any restrictions on where it can reside. If moving the data adds more infrastructure to manage without a meaningful performance benefit, keeping processing closer to the source may make more sense.
This is not a blanket “never move data” strategy. It is a workload-based approach: keep data close to its source when locality, governance, freshness, or scale matter, and move or copy data when the workload has a clear reason to do so.
| Consideration | Keep processing near the data when … | Move data when … |
| Privacy or residency | Data must remain in a controlled location | Policy allows transfer to another environment |
| Data volume | Transfers would be costly or slow | The dataset is manageable to move |
| Freshness and latency | The workload needs current, low-latency access | Some delay is acceptable |
| Compute location | Suitable processing resources are already nearby | Specialized compute is available elsewhere |
| Multiple sources | Most required data is already local | The workload must combine data from several locations |
As enterprise data spreads beyond traditional data centers, organizations may need to process more of it closer to where it is generated. Keeping processing closer to the data can reduce transfers between centralized locations while supporting security, audit, and data sovereignty requirements.
AI workloads can also run where the data already resides. The Dell AI Data Platform with NVIDIA is the data foundation of the Dell AI Factory with NVIDIA, helping organizations connect governed enterprise data to AI workloads while keeping data, models, and operations within trusted environments. Organizations can use the approach when privacy, residency, or the volume of data makes large transfers impractical.
According to NVIDIA, BlueField-3 data processing units (DPUs) accelerate access across storage initiators, controllers, and targets, allowing multiple GPUs to share the same data stores. NVIDIA says this approach can reduce CPU overhead as AI workloads scale.
How can companies improve data accessibility for AI while maintaining governance and control?
Companies can improve data accessibility for AI by giving approved workloads access to the data they need while enforcing governance controls. Existing ownership and access rules still need to apply as more models and applications connect to enterprise data.
Practical governance controls can include:
- Access controls: Limit data access by role, attribute, or workload.
- Ownership and lineage: Record who owns the data and how it changes over time.
- Sensitive-data protection: Use masking, redaction, or tokenization where needed.
- Retention rules: Define when data should be retained or deleted.
- Ongoing monitoring: Track data quality and access over time.
Governed pipelines can apply these controls during data preparation. Dell AI Data Platform with NVIDIA can index and prepare unstructured information and connect it to governed pipelines for data discovery and dataset creation. A managed path for preparation can reduce reliance on ad hoc copies and give teams a defined way to apply access, lineage, quality, and governance requirements before data reaches an AI workload.
NVIDIA-accelerated infrastructure and AI software contribute security controls to the integrated solution, including policy enforcement, isolation, and encryption for data at rest and in transit. These controls complement Dell’s storage, cyber resilience, access, and governance capabilities.
Governance also depends on knowing who owns the data, who can use it, and how it is managed over time. As part of its developing Data Governance and Management Profile, the National Institute of Standards and Technology (NIST) emphasizes managing access and maintaining data quality. It also calls for enough metadata to show where data came from and how it has changed over time. Clear ownership and consistent access rules give teams a way to make data available to AI applications without losing oversight of how it is used.
How do data pipelines and orchestration help companies scale AI?
Data pipelines and orchestration help companies scale AI by making preparation and delivery repeatable. Gartner recommends preparing data pipelines for both model-training datasets and live data feeds to production AI systems.
Data pipelines handle recurring preparation work, while data orchestration coordinates when those pipelines run and how information moves between stages. Orchestration can also manage scheduling, retries, and task dependencies while tracking how data moves through the pipeline. Together, they give teams a more consistent process for applying validation and governance as new AI workloads move into production.
The Dell AI Data Platform with NVIDIA can index billions of unstructured files and connect them into governed pipelines, speeding data discovery and dataset creation. Dell Technologies' data services address data preparation and operational complexity as AI projects move from pilots into production.
More than 6,5000 customers were deploying Dell AI Factory as of August 2026, based on Dell's analysis of customer order data. The Dell AI Factory with NVIDIA supports AI workloads from training through deployment, combining Dell infrastructure and services with NVIDIA accelerated computing and AI software. A strong enterprise AI data strategy starts with a simple question: Where should each workload access and process the data it needs? Answering that question well can reduce unnecessary movement while giving enterprises more control as AI projects scale.