# Storage Systems > [!abstract] The idea > The answer to "where should I store this data?" is not always a database. Different workloads need different storage *models*, and the right one is chosen by **how the data will be accessed, how fast it will grow, and what guarantees it needs**, not by where there happens to be space. Part of [[Storage Systems MOC]] · Related maps: [[085 Systems MOC]] · [[Data Center MoC]] · [[Open Source Hyperscaler MoC]] ![[Storage Systems - Infographic.jpg|700]] ## The core move: reframe the question The weak question is *"Where can I store this data?"* The strong question is: 1. **How will it be accessed?** Random small reads, sequential streams, whole-object fetches, shared by many clients? 2. **How much will it grow?** Gigabytes that fit on one disk, or petabytes that need to scale out? 3. **What guarantees does it need?** Latency, durability, availability, consistency, jurisdiction. Everything else follows. Storage is a trade between those answers and cost, and no single system wins all of them. This is the same discipline as in [[System Design in the Age of AI]]: name what can't fail, then let the rest fall out. ## Three models, three contracts | Model | What you get | Access contract | Natural home | | --- | --- | --- | --- | | **Block** | Raw volumes, used like a disk | Lowest latency, random read/write at fixed-size blocks | Databases, VM disks | | **File** | Files and folders with shared access | Hierarchical paths over the network (NFS/SMB), many clients | Shared files, enterprise and HPC workloads | | **Object** | Objects addressed by key through an API | Whole-object get/put, flat namespace, near-unbounded scale | Images, video, backups, logs, large datasets | The deeper pattern is a **ladder of abstraction traded against scale**. Block gives you the least structure and the most speed. File adds structure (hierarchy, shared semantics) and pays for it in coordination. Object drops the hierarchy and POSIX semantics entirely, and in return scales horizontally almost without limit and costs the least per byte. > [!tip] Rule of thumb > The more a workload needs *mutation in place at low latency*, the further toward block it sits. The more it is *write once, read many, very large*, the further toward object. ## Worked example: why video never goes in the database ```mermaid graph LR U[User uploads 2 GB video] --> A[Application Server] A --> O[(Object Storage)] O --> C[CDN] C --> D[User downloads video] A -. metadata only .-> DB[(Database)] ``` - The **database holds metadata** (who, what, where, permissions) because that is small, relational and mutable. - **Object storage holds the bytes** because that is large, immutable and read-heavy. - The **CDN** moves the read path to the edge so the origin is not the bottleneck. Putting a 2 GB blob in a database couples the cheap-to-scale thing (bytes) to the expensive-to-scale thing (transactional state). Separating them lets each scale on its own curve. ## Workload to storage mapping - Database or VM disk → **block** (low latency, high IOPS) - Shared files across clients → **file** (shared access semantics) - Images, video, backups, static assets → **object** (cheap, durable, scalable) - Application logs → **object** (cheap, append and retain) - Large analytic datasets → **object / distributed storage** (massive volume) ## The seven forces Every storage decision is a weighting of these. Decide which one is the non-negotiable first. - **Latency**: how fast a single read or write returns - **Throughput**: how much data moves per second in aggregate - **Scalability**: whether it grows by adding nodes or by buying bigger boxes - **Durability**: whether data survives failures (replication, erasure coding) - **Availability**: whether data is reachable right now, a separate property from durability - **Cost**: price per byte and per operation at scale - **Access pattern**: random vs sequential, hot vs cold, read vs write heavy > [!note] Durability is not availability > A system can lose none of your data and still be down. Eleven nines of durability says nothing about uptime. Design for each separately, see [[Data Centre Redundancy]] and [[Data Center Tiers]] for the physical analogue. ## The memory hierarchy is the same idea, zoomed in The storage choice at the system level mirrors the memory choice at the chip level: fast, small and expensive on one end, slow, vast and cheap on the other. `SRAM → HBM → DRAM → NVMe (block) → HDD / object → cold archive` The same trade-offs reappear whether you are choosing between [[HBM]] and [[GDDR]], or between block and object. See [[Memory Bandwidth]], [[SRAM]] and [[KV Cache Compression and Eviction]] for the AI-inference end of the ladder, and [[Cache Coherent Interconnect (CXL)]] for the attempt to blur the boundary between memory and storage. ## Disaggregation: storage as a pool, not a box The long-run direction is **compute and storage becoming independent pools** joined by a fast fabric, instead of disks bolted inside servers. [[NVMe Fabric]] and [[SPDK]] are how block-level latency survives being moved across the network, and it is the storage-side version of the shift described in [[Data Center MoC]] (monolithic server → composable → disaggregated). The consequence: capacity and performance can be scaled and purchased separately, which is also a utilisation argument (see [[Capacity Managament]] and [[Puzzle of low data center utilisation]]). ## One system, three faces [[Ceph]] is the clearest proof that the three models are interfaces rather than separate technologies. One distributed object store (RADOS) underneath, exposed as block (RBD), object (RadosGW) and file (CephFS). The unification is at the substrate, the distinction is at the contract. Contrast with specialising: [[MinIO]] does only object and wins on simplicity and throughput, [[Longhorn]] does only Kubernetes-native block and wins on operational ease. The decision is *breadth of one platform* vs *excellence per workload*. ## Object storage is the substrate for everything else A lot of the observability and analytics world quietly sits on object storage because it makes retention a storage-price problem rather than a compute-price problem: - [[Loki]], [[Thanos]], [[Mimir]], [[Tempo]] all ship their long-term data to S3-compatible storage - [[data lakes]] are literally compute pointed at cheap blob storage - Backups and snapshots for block volumes (e.g. [[Longhorn]]) land on object storage The pattern: **hot state on block, durable bulk on object, compute on top of both.** ## Storage as a sovereignty question Once data has a jurisdiction, "where" returns, but as a *legal and cryptographic* property rather than a disk location. Storage becomes a question of who can compel access and who holds the keys. - [[Navon Sovereign Vaults]] and [[Data Embassies]] treat storage as a jurisdictional guarantee, with [[Hybrid Cryptography as the bridge to Post Quantum Security]] protecting it over long horizons - [[Compute to data]] and [[Confidential Computing]] invert the flow: instead of moving data to compute, move compute to where the data must stay - [[Nirvana_Africa_Sovereign_Storage_Strategy.docx]] applies this to the African context High-performance storage at the edge on stable power is a recurring part of the [[Navon Thesis]] and is the practical bridge between this note and the infrastructure thesis. ## Storage for agents Agents are stateful by definition, so the storage layer decides what they can remember. See [[AI Agents Stack]] and [[Persistent Memory]]: vector stores and retrieval give recall, not continuity, which is the same block / object / hot / cold split replayed at the application layer. The [[Latest Stack]] covers the concrete choices. ## Takeaways - Match the **storage model to the access pattern**, not to habit. - Keep **metadata in the database and bytes in object storage**. - Separate **durability from availability** and decide each on purpose. - The block / file / object split is a **contract**, and one platform can offer all three. - The trend is **disaggregation**: storage scales as its own pool. - Ask: *how will this data be accessed, how much will it grow, what guarantees do I need?* > [!question] Open threads > - Where does the line fall between "distributed file system" and "object storage with a POSIX gateway" for AI training data? > - How much of sovereignty is storage design vs key custody? #system-design #storage #firstprinciple