The technical anatomy of early Google search
Summary
Sergey Brin and Lawrence Page describe a search engine that combines link structure, anchor texts and scalable data processing for more relevant results. The prototype described indexed about 24 million web pages. PageRank was calculated iteratively from the link structure of the web.
Ideas
- PageRank uses links as weighted recommendations between web pages.
- Anchor texts describe target pages even when their own text is insufficient.
- A URL server coordinates which addresses the distributed crawler fetches next.
- Barrels store sorted word hits for fast queries in the inverted index.
- Position in the document and font features influence the weighting of a hit.
- A clear data pipeline separates crawling, analysis, indexing and query processing.
Insights
- Relationships between documents can carry more meaning than document content alone.
- Useful scaling starts with suitable data formats and clear processing steps.
- Quality emerges from several weak signals evaluated together.
- A research system gains credibility through real size and measurable performance.
Facts
- The system stored anchor text together with information about its link target.
- The architecture ran on inexpensive commodity hardware under Linux.
Recommendations
- Clearly separate data collection, normalisation, index building and querying.
- Use external relationships as a relevance signal of their own.
- Store intermediate formats so that every processing step remains repeatable.
References
Links to the original source and the Web Archive open in a new tab.