Migration to premium SSD: query latency halved day one. Three weeks later, slow again. VPS volume burst credits permanently exhausted — invisible 300 IOPS cap on pricing page.
IOPS explains fast-at-first vs sustainable 24/7 disk.
Random vs sequential
Marketing cites sequential GB/s. Your DB does random 4K. Measure what matters.
Burst and credits
Cheap VPS/AWS burst then throttle. iostat saturation signals. Read baseline vs burst contract.
App impact
Postgres checkpoints, MySQL flush, Redis AOF. CPU idle, app slow, wa high.
Sizing
Estimate needed IOPS + 30%. Separate data/WAL. Object storage for blobs.
Compare hosts
Identical sustained 30-min benchmark, not 30 s.
Minimal benchmark
fio 4k randrw 70/30, iodepth 32, runtime 600s, direct=1. Note avg IOPS and p99 latency.
Repeat under app load if possible.
Compare three providers same price tier — 2× gap common.
Prod DB: separate data and WAL volumes if provider allows.
Renegotiate with numbers, not "it feels slow".
App tuning before disk upgrade
Postgres: checkpoint tuning, SSD random_page_cost. MySQL: innodb flush log at transaction commit vs 2 — durability/IOPS tradeoff.
Disk upgrade without app tuning = repeated disappointment.
Elasticsearch on network disk without provisioned IOPS: permanent yellow cluster — size correctly or local SSD.
Small random fsync-heavy writes: NVMe marketing ≠ fsync latency under load.
Cloud volume types
gp3 IOPS provisioned independent of size — upgrade IOPS without GB resize.
Local NVMe instance store ephemeral: high IOPS but loss on instance stop — WAL on network block with IOPS.
Monitoring: iostat await p99 alert >20ms sustained.
Operational summary
Sustained mixed read/write IOPS determine DB feel — not NVMe marketing peak. fio 600s, iostat await, separate WAL/data if possible. Disk upgrade without checkpoint tuning = repeated disappointment.
Compare hosts with same benchmark script before signing.
Before premium disk purchase
Archive current fio baseline. Upgrade tier. Re-fio 600s same script. Compare % gain vs +30% cost. If gain <15%: app tuning first.
Document finance decision memo — avoids annual disk upgrade cycle without proof.
Provider negotiation
Attach benchmark PDF to renewal — providers sometimes match competitor provisioned IOPS without tier jump if evidence clear.
Local SSD instance store for cache layer plus network block for durability splits cost and performance sensibly.
Index review
Before disk tier upgrade, run slow query log week and explain analyze top ten. Missing index fix often drops IOPS need 50%+. Share explain results in infrastructure ticket — proves engineering due diligence.
Read replica for reporting offloads primary IOPS — analytics queries not on OLTP primary.
Cloud volume resize online still needs filesystem grow inside OS — runbook both steps.
Operational follow-up
IOPS graph in monthly exec dashboard — budget visibility. Correlate cloud bill and provisioned IOPS. Postgres pg_stat_statements top ten monthly — index before invoice argument. Document gaps between host marketing and field measurement in the quarterly review.
Quarterly follow-up
IOPS graph in monthly exec dashboard — budget visibility. Correlate cloud bill and provisioned IOPS. Postgres pg_stat_statements top ten monthly — index before invoice argument. Document gaps between host marketing and field measurement in the quarterly review.
Keep a dated runbook, before/after metrics, post-incident review — cumulative discipline beats Friday night panic.
Keep a dated runbook, before/after metrics, post-incident review — cumulative discipline beats Friday night panic.
Keep a dated runbook, before/after metrics, post-incident review — cumulative discipline beats Friday night panic.
Keep a dated runbook, before/after metrics, post-incident review — cumulative discipline beats Friday night panic.
Keep a dated runbook, before/after metrics, post-incident review — cumulative discipline beats Friday night panic.
Keep a dated runbook, before/after metrics, post-incident review — cumulative discipline beats Friday night panic.
fio and tuning
600s fio before premium tier. <15% gain → Postgres tuning first. Renewal PDF. Data/WAL split.
Renegotiation
Attach identical benchmark three providers same tier — 2× gap common.
Decide and move forward without blind spots
- Identical fio benchmark — 4k randrw 70/30, iodepth 32, 600s, direct=1 on idle volume then under app load.
- Compare three providers same price tier — 2× gap common; archive PDF for renewal.
- Separate data and WAL — if provider allows on prod PostgreSQL.
- Read sustained mixed IOPS — not ten-second marketing burst.
- Renegotiate with numbers — not « it feels slow » without iowait and queue depth.
Pick storage tier via our directory, comparison tool, and database guides.
Frequently asked questions
IOPS vs MB/s throughput?
IOPS = 4K random ops/s. MB/s = sequential. OLTP DB = random IOPS; backup/video = throughput.
Why unrealistic marketing numbers?
Short burst, ideal 4K read, local cache — prod mixed sustained is much lower.
How to test?
fio randrw 4k on idle volume AND under app load. Compare providers.
Undersizing symptoms?
High iowait, long Postgres checkpoints, MySQL commit latency, disk queue depth > 10 sustained.
Ask support for sustained IOPS cap before next Black Friday — burst does not cover eight-hour peaks.
