Breaking
Smart Devices

AI data boom reshapes storage priorities over compute

By Blake Weston 4 min read
AI data boom reshapes storage priorities over compute - ai storage priorities
IDC’s white paper Built for Scale examines AI data persistence outlasting training cycles by years.

AI’s data explosion is forcing companies to rethink storage—not just compute. A new global study from IDC, sponsored by WD, shows AI workloads are creating a structural shift in storage needs, with data growing faster than compute demands and persisting long after training cycles end.

The focus on AI infrastructure has been dominated by compute power—GPUs, TPUs, and cloud clusters—but the IDC white paper, Built for Scale: The Enduring Role of HDDs in the AI Era, highlights a less-discussed but equally critical challenge: data that doesn’t disappear after use. Unlike compute cycles, which can be scaled up or down, the data AI generates, training logs, synthetic datasets, inference outputs, accumulates and retains value over time.

Among the surveyed organizations, 94.7% reported storing more data in the past year due to AI and generative AI adoption. 61% saw data volumes grow by 25% or more over the same period, while 74% expect similar growth in the next three years. The expansion isn’t just about raw volume: 85.4% also reported growth in data lake volumes, with 59.4% citing AI-generated data, including synthetic data and model logs, as the primary driver.

The study shows a fundamental shift in how companies treat data. Nearly 95% say the value of their data has increased because of AI, and 74.3% are retaining data longer than before. Even archived “cold” data, once considered dead storage, is being reactivated. 75.9% now bring historical datasets back online to fuel AI workloads, and 96% anticipate needing faster retrieval from archives to support applications like retrieval-augmented generation (RAG).

Storage costs now dictate AI strategy

This isn’t just about capacity. The economics of storage are becoming a deciding factor in AI strategy. 98.2% of organizations rank total cost of ownership per terabyte as important or very important when selecting storage solutions. The implication is clear: AI success now depends on managing data across its entire lifecycle, from active workloads to long-term archives, rather than treating storage as an afterthought.

Read Also: Redington partners AvePoint to enhance APAC data resilience

The traditional divide between active and inactive data is blurring. Historical datasets, once relegated to cheap, slow storage, are now being repurposed for fine-tuning models, generating synthetic data, or enriching training sets. The result is a compounding effect: AI creates data, that data gains value over time, and organizations hold onto it longer, repeating the cycle. This forces infrastructure teams to balance performance, accessibility, and cost in ways they didn’t before.

For companies, this means storage isn’t just a supporting function, it’s a competitive advantage. The ability to retain, access, and repurpose data at scale could determine which organizations lead in AI innovation. Yet the study also reveals a gap: 74.6% of enterprise data already lives in warm, cool, or cold storage tiers, with over 60% of data lake volumes consisting of cold or infrequently accessed data. The challenge isn’t just storing more; it’s architecting systems that can efficiently tier data based on usage patterns and cost.

Data persistence becomes AI’s hidden advantage

WD’s CEO, Irving Tan, framed it simply: “For the last few years, the AI infrastructure conversation has centered on compute. But AI runs on data.” The shift isn’t just about buying more storage. It’s about designing systems where data remains useful, accessible, and economical across its entire lifespan. Compute will always matter, but the study suggests that without a storage strategy built for persistence and reuse, even the most powerful AI models will hit a wall.

The findings also hint at a broader industry reckoning. Storage vendors, cloud providers, and enterprises are all adapting to this new reality. The question now isn’t whether data will keep growing, it’s how organizations will structure their infrastructure to handle it without breaking the bank. For now, the answer lies in balancing high-performance storage for active workloads with cost-effective solutions for the long tail of data that never truly goes away.

The study doesn’t address whether current storage technologies can keep up. But one thing is clear: the days of treating storage as a secondary concern in AI infrastructure are over. The data isn’t going away, and neither is its value.

Blake Weston

Leave a Reply

Your email address will not be published. Required fields are marked *