Load balancing
A load balancer directs incoming requests to servers that are available, rather than letting every request pile onto one machine.
Google, YouTube, Amazon, and Facebook do not rely on one enormous computer. They divide the work among many systems, each designed to do a particular job well.
When you open a page, your browser asks for information. A server receives that request, runs the application code it needs, looks up any required data, and sends a response back.
For a small site, one server can comfortably handle that work. The problem begins when too many people ask at the same time: every computer has limits on processing, memory, storage, and network capacity.
Large services scale by spreading work across groups of computers. More visitors do not need a magical supercomputer; they need enough capacity, sensible routing, and systems that can work together.
A load balancer directs incoming requests to servers that are available, rather than letting every request pile onto one machine.
Frequently requested information can be kept close at hand, so the same work does not have to be repeated for every visitor.
When demand grows, a service can add more application servers, storage, and network capacity to handle it.
Large services hold far more than web pages: accounts, messages, search indexes, photographs, videos, and the records that connect all of it. That information is spread across storage and database systems rather than kept on one drive.
Copies also matter. If an important system or storage device fails, another copy can keep the service available. Managing those copies is difficult, especially when information changes, but it is a major part of making a large service dependable.
Software tracks where data belongs and retrieves it when needed, while many storage systems share the volume of files and records.
Important data can exist in more than one place so a single hardware failure does not make it disappear.
Authentication, search, databases, queues, and media delivery may all be separate systems, tuned for the work they do.
People use global services from all over the world. The distance between a visitor and the systems serving them can affect speed, so large companies operate in multiple regions and use caching and content-delivery networks to move common content closer.
No system can assume every computer, cable, or data center will work forever. Large services plan for individual failures and, where practical, larger outages by keeping spare capacity and independent copies in different places.
The closest server is not always the right answer. A cached image may come from nearby, while an account change may need to reach the specific database responsible for that information. The systems decide where each kind of work should happen.
To you, it is a search, a click, or a video that starts playing. Behind that moment, many systems may be sharing requests, storing data, handling failures, and keeping common content ready. The real achievement is making all of that feel like one dependable service.