Overview
Amazon S3, launched in 2006 by AWS, is a scalable and durable cloud storage service for unstructured data. It’s ideal for use cases like data lakes, backups, and disaster recovery. It offers storage for objects up to 5 TB, with global access.
Key Features
- Scalable: Automatically adjusts to storage needs.
Storage Classes
- S3 Standard: Frequently accessed data.
- S3 Standard-IA: Infrequent access with lower costs.
- S3 Glacier: Archival storage with flexible retrieval.
- S3 Glacier Deep Archive: Cheapest for long-term storage.
How S3 Works
S3 stores data as objects in buckets. Objects include data, metadata, and a unique identifier. Data is replicated across availability zones for durability.
Core Operations
- Bucket Creation: Managed via API, metadata stored in database.
- Object Download: Retrieved through metadata service.
Low-Level Design
- Data Service: Manages storage nodes and data replication.
- Metadata Service: Manages object and bucket metadata.
- WAL: Objects are written to logs for efficiency.
Expansion
Data Integrity
Uses checksums to detect corruption.
Versioning
Supports multiple versions of objects.
Multipart Uploads
Efficient large file uploads.
Amazon S3 utilizes object storage technology, emphasizing its durability, availability, and performance. It has become a standard in the industry, widely adopted not just by Amazon Web Services (AWS), but also by providers like Cloudflare, Backblaze, and Google Cloud Storage.
Scalability and Availability
Amazon S3 is highly scalable, providing 99.99% availability, ensuring your data is always accessible, and designed for 99.999999999% (11 nines) durability, assuring that your data is protected.
Storage Class Options for Infrequently Accessed Data
- S3 Standard-IA: Lower storage costs than S3 Standard with higher retrieval costs.
Use Case: Backup and disaster recovery, archival data that is infrequently accessed. - S3 One Zone-IA: Similar to Standard-IA but stores data in a single availability zone for lower costs.
Use Case: Non-critical, easily reproducible data.
Storage Class Options for Frequently Accessed Data
- S3 Glacier: Low cost of $0.004 per GB-month for data that is rarely accessed but needs immediate retrieval.
Use Case: Archived video footage requiring fast access. - S3 Glacier Deep Archive: Most cost-effective for long-term archival storage.
Use Case: Old log files for regulatory compliance.
How Amazon S3 Works: Understanding Cloud Object Storage
Amazon S3 is designed around the concept of object storage, storing data as objects within buckets, where each object consists of the actual data, metadata, and a unique identifier. S3 replicates objects across multiple availability zones to ensure data durability and availability.
Data Service and Metadata Service
- Data Service: Reading and writing objects to disk.
- Metadata Service: Handles metadata, including object names, IDs, and storage locations.
Core Operations for S3
- Bucket Creation: An HTTP PUT request is sent to the API service. The metadata service stores the bucket information.
- Object Upload: An HTTP PUT request stores the object, checked for authentication and permissions. The metadata service stores the object’s metadata.
- Object Retrieval: An HTTP GET request retrieves the object after verifying permissions through the metadata service.
System Expansion
Amazon S3 manages vast amounts of data with high availability and scalability, making it ideal for many applications. Efficient management of storage nodes and intelligent data replication strategies are critical for optimal performance and cost-effectiveness, emphasizing features like data integrity through heartbeats and replication across nodes.