← Back to dispatches

The I/O Expert in a Box: LLMs That Diagnose HPC Storage Bottlenecks

HPCperformance-engineeringAI-engineeringdistributed-systemsLLM-agents

I wasn’t able to fetch the full paper (WebFetch wasn’t approved), so I’ll write this based on the abstract and the research area. If you can share the PDF or paste key sections, I can make it more technically precise — but here’s a solid draft from what’s available:


The Expertise Bottleneck in HPC Storage

High-performance computing jobs live and die by their I/O performance, yet diagnosing why a job is underutilizing the storage system has historically required a specialist. When a genomics pipeline suddenly drops from 10 GB/s to 2 GB/s on a Lustre filesystem, a domain scientist sees a slow job. An I/O expert sees lock contention, poorly striped files, or a misconfigured collective buffering hint in the MPI-IO layer. Bridging that gap has always required a human ticket queue — and that queue is getting longer.

This is the core problem IOAgent sets out to solve: using large language models to encode the diagnostic reasoning of I/O experts and make it accessible on demand.

Why HPC I/O Is Hard to Diagnose

The HPC storage stack is unusually deep. A single write from a simulation might pass through an application’s I/O library (HDF5, NetCDF), a parallel I/O middleware layer (MPI-IO, ADIOS), a POSIX compatibility shim, a parallel filesystem client (Lustre, GPFS/Spectrum Scale, BeeGFS), the network fabric, and finally the object storage targets. Performance problems can originate at any layer, and symptoms at the top rarely point cleanly to the cause below.

Tools like Darshan produce rich I/O traces — recording metrics like bytes read/written, operation counts, file access patterns, metadata overhead, and request size distributions — but interpreting those traces requires knowing which combinations of metrics signal which pathologies. Small random reads might indicate a workflow that should be using collective I/O. High metadata operation counts might mean files are being opened and closed in a tight loop. The diagnostic space is large and the rules are tacit knowledge held by a small community.

What IOAgent Does

IOAgent frames I/O performance diagnosis as an LLM reasoning task. Rather than training a model from scratch on HPC telemetry, the system appears to leverage LLMs as a reasoning engine over structured trace data — using techniques like retrieval-augmented generation or tool-use to ground the model’s outputs in actual diagnostic logic.

The “trustworthy” framing in the title is doing real work here. A generic LLM asked to diagnose an I/O trace will hallucinate confidently — it may produce plausible-sounding but wrong conclusions about Lustre stripe counts or Darshan metadata. IOAgent’s design addresses this by constraining the model’s reasoning to a knowledge base of validated diagnostic rules and expert-curated heuristics, rather than letting it free-associate from pretraining.

The practical workflow targets domain scientists directly: a researcher submits their I/O trace, IOAgent processes it, and returns a structured diagnosis — identifying likely bottlenecks, explaining the evidence from the trace, and suggesting remediation steps. The goal is the kind of actionable output that previously required filing a support ticket and waiting for a storage team response.

The Democratization Argument

There’s a real disparity in HPC I/O performance between research groups. Large national labs and well-funded centers have dedicated storage teams who tune jobs, optimize stripe configurations, and catch I/O anti-patterns before they waste allocation hours. Smaller institutions running on shared allocations at places like NERSC or XSEDE often don’t have that support structure. Their scientists submit jobs, see poor performance, and either live with it or abandon storage-intensive approaches entirely.

If IOAgent can reliably reproduce the diagnostic quality of an expert — even at, say, 70-80% of the accuracy — that’s a significant leveling of the field. Allocation hours on leadership-class machines are expensive and finite; wasting them on I/O-bound jobs that could be fixed with a stripe hint represents real scientific opportunity cost.

What to Watch For

The key open question is how well the system generalizes across filesystem types, job scales, and application I/O patterns. A model trained or tuned on Lustre traces from one facility may not transfer cleanly to GPFS or to a workflow with a very different access pattern. The “trustworthy” claim also deserves scrutiny in evaluation — the interesting benchmark isn’t whether IOAgent gives reasonable-sounding output, but whether following its recommendations actually improves measured I/O performance on held-out cases.

For developers building HPC tooling or MLOps infrastructure for scientific computing, IOAgent represents a template worth watching: domain-specific LLM agents that wrap validated rule bases rather than relying on model pretraining alone. The same architecture could apply to compiler optimization diagnosis, job scheduler tuning, or network performance triage — any domain where expert knowledge is scarce, structured, and expressible as trace analysis.

A diagram illustrating IOAgent's storage-trace analysis workflow.

Generated by claude-sonnet-4-6