Storage System 101
In this chapter, we design an object storage service similar to Amazon Simple Storage Service (S3). S3 is a service offered by Amazon Web Services (AWS) that provides object storage through a RESTful API-based interface. Here are some facts about AWS S3:
-
Launched in June 2006.
-
S3 added versioning, bucket policy, and multipart upload support in 2010.
-
S3 added server-side encryption, multi-object delete, and object expiration in 2011.
-
Amazon reported 2 trillion objects stored in S3 by 2013.
-
Life cycle policy, event notification, and cross-region replication support were introduced in 2014 and 2015.
-
Amazon reported over 100 trillion objects stored in S3 by 2021.
Before we dig into object storage, let’s first review storage systems in general and define some terminologies.
Storage System 101
At a high-level, storage systems fall into three broad categories:
-
Block storage
-
File storage
-
Object storage
Block storage
Block storage came first, in the 1960s. Common storage devices like hard disk drives (HDD) and solid-state drives (SSD) that are physically attached to servers are all considered as block storage.
Block storage presents the raw blocks to the server as a volume. This is the most flexible and versatile form of storage. The server can format the raw blocks and use them as a file system, or it can hand control of those blocks to an application. Some applications like a database or a virtual machine engine manage these blocks directly in order to squeeze every drop of performance out of them.
Block storage is not limited to physically attached storage. Block storage could be connected to a server over a high-speed network or over industry-standard connectivity protocols like Fibre Channel (FC) 1 and iSCSI 2. Conceptually, the network-attached block storage still presents raw blocks. To the servers, it works the same as physically attached block storage.
File storage
File storage is built on top of block storage. It provides a higher-level abstraction to make it easier to handle files and directories. Data is stored as files under a hierarchical directory structure. File storage is the most common general-purpose storage solution. File storage could be made accessible by a large number of servers using common file-level network protocols like SMB/CIFS 3 and NFS 4. The servers accessing file storage do not need to deal with the complexity of managing the blocks, formatting volume, etc. The simplicity of file storage makes it a great solution for sharing a large number of files and folders within an organization.
Object storage
Object storage is new. It makes a very deliberate tradeoff to sacrifice performance for high durability, vast scale, and low cost. It targets relatively “cold” data and is mainly used for archival and backup. Object storage stores all data as objects in a flat structure. There is no hierarchical directory structure. Data access is normally provided via a RESTful API. It is relatively slow compared to other storage types. Most public cloud service providers have an object storage offering, such as AWS S3, Google object storage, and Azure blob storage.
Comparison
Table 1 compares block storage, file storage, and object storage.
| Block storage | File storage | Object storage | |
|---|---|---|---|
| Mutable Content | Y | Y | N (object versioning is supported, in-place update is not) |
| Cost | High | Medium to high | Low |
| Performance | Medium to high, very high | Medium to high | Low to medium |
| Consistency | Strong consistency | Strong consistency | Strong consistency [5] |
| Data access | SAS [6]/iSCSI/FC | Standard file access, CIFS/SMB, and NFS | RESTful API |
| Scalability | Medium scalability | High scalability | Vast scalability |
| Good for | Virtual machines (VM), high-performance applications like database | General-purpose file system access | Binary data, unstructured data |
Table 1 Storage options
Terminology
To design S3-like object storage, we need to understand some core object storage concepts first. This section provides an overview of the terms that apply to object storage.
Bucket. A logical container for objects. The bucket name is globally unique. To upload data to S3, we must first create a bucket.
Object. An object is an individual piece of data we store in a bucket. It contains object data (also called payload) and metadata. Object data can be any sequence of bytes we want to store. The metadata is a set of name-value pairs that describe the object.
Versioning. A feature that keeps multiple variants of an object in the same bucket. It is enabled at bucket-level. This feature enables users to recover objects that are deleted or overwritten by accident.
Uniform Resource Identifier (URI). The object storage provides RESTful APIs to access its resources, namely, buckets and objects. Each resource is uniquely identified by its URI.
Service-level agreement (SLA). A service-level agreement is a contract between a service provider and a client. For example, the Amazon S3 Standard-Infrequent Access storage class provides the following SLA 7:
-
Designed for durability of 99.999999999% of objects across multiple Availability Zones.
-
Data is resilient in the event of one entire Availability Zone destruction.
-
Designed for 99.9% availability.
Finished reading?
Mark it complete to track your progress.