Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

WebArchive — Self-Hosted Wayback Machine

A lightweight, self-hosted web archiving system. Archive any URL, subdomain, or entire domain — snapshots stored locally, viewable offline.

Features

  • Single URL, multi-URL, subdomain, or full-domain crawling
  • One snapshot per URL — archiving again overwrites the previous
  • Content-addressed asset storage — identical assets stored once via SHA-256 hash
  • Playwright-powered crawling — handles JS-rendered pages
  • Snapshot viewer — view archived pages in-browser with local assets
  • Real-time job monitoring — live progress updates for active crawls
  • SQLite database — zero-dependency metadata storage
  • Fully Dockerized — one command to run

Quick Start

Requirements

  • Docker & Docker Compose

Run

git clone https://github.com/nr-yolo/Web-Archiver.git
cd webarchive
mkdir data
docker compose up -d --build

Open http://localhost:8777 in your browser.


Architecture

┌─────────────────┐     ┌──────────────────────┐
│  React Frontend │────▶│  FastAPI Backend      │
│  (Nginx :8080)  │     │  (Uvicorn :8000)      │
└─────────────────┘     │                       │
                        │  ┌─────────────────┐  │
                        │  │ Playwright       │  │
                        │  │ Crawler Engine  │  │
                        │  └────────┬────────┘  │
                        │           │            │
                        │  ┌────────▼────────┐  │
                        │  │ SQLite Database │  │
                        │  │ /data/webarchive│  │
                        │  │   .db           │  │
                        │  └────────┬────────┘  │
                        │           │            │
                        │  ┌────────▼────────┐  │
                        │  │ File Storage    │  │
                        │  │ /data/snapshots │  │
                        │  └─────────────────┘  │
                        └──────────────────────┘

API Reference

All endpoints are under /api/. No authentication required.

POST /api/archive

Start an archiving job.

{
  "urls": ["https://example.com"],
  "scope": "single",
  "max_depth": 3,
  "max_pages": 100
}

Scopes:

  • single — archive only the submitted URLs, no crawling
  • subdomain — crawl all pages within the same subdomain
  • domain — crawl the entire domain

Response:

{
  "job_id": "uuid",
  "status": "pending",
  "urls": ["https://example.com"]
}

GET /api/archives

List all archived snapshots.

Query params: q (search), limit, offset


GET /api/snapshot/{url}

View a snapshot. {url} should be URL-encoded.


GET /api/asset/{content_hash}

Serve a stored asset (CSS, JS, image, font, etc.)


DELETE /api/snapshot/{url}

Delete a snapshot and its files.


GET /api/jobs

List recent crawl jobs.

GET /api/jobs/{job_id}

Get status of a specific job.

GET /api/stats

System statistics (snapshot count, asset count, storage size).


Storage Layout

/data/
├── webarchive.db           # SQLite metadata
└── snapshots/
    ├── _assets/            # Content-addressed assets
    │   ├── ab/
    │   │   └── ab3f9c....css
    │   └── ...
    ├── https___example_com/
    │   └── index.html      # Archived page (assets rewritten)
    └── https___another_com_page/
        └── index.html

Crawler Behavior

Scope enforcement

  • Single: fetches only the given URL(s), no link following
  • Subdomain: follows links within the same host (e.g., blog.example.com)
  • Domain: follows links within the same registered domain (e.g., *.example.com)

External resources

External CSS, JS, fonts, and images are fetched and stored locally (needed to render pages correctly). External HTML pages are not followed.

Asset deduplication

All assets are stored by SHA-256 hash. If two pages use the same CSS file, it's stored once. The HTML is rewritten to point to /api/asset/{hash}.

Internal link rewriting

Links between archived pages are rewritten to /snapshot/{url} so navigation stays within the archive.


Configuration

Environment variables (set in docker-compose.yml):

Variable Default Description
DB_PATH /data/webarchive.db SQLite database location
STORAGE_PATH /data/snapshots Snapshot storage root

Development

Backend (FastAPI)

cd backend
pip install -r requirements.txt
playwright install chromium
uvicorn main:app --reload --port 8000

Frontend (React)

cd frontend
npm install
REACT_APP_API_URL=http://localhost:8000/api npm start

Tech Stack

Layer Technology
Frontend React 18, React Router, react-hot-toast
Backend FastAPI, Uvicorn
Crawler Playwright (Chromium)
HTTP aiohttp, httpx
HTML parsing BeautifulSoup4, lxml
Database SQLite via aiosqlite
Storage Local filesystem
Serving Nginx (frontend), Uvicorn (backend)
Container Docker, Docker Compose

Notes

  • Crawling large sites can be slow — Playwright launches a full Chromium instance per crawl job
  • For very large sites (1000+ pages), consider increasing max_pages carefully
  • The archive viewer opens snapshots in a new tab; internal links navigate between archived pages
  • No authentication — intended for local/trusted network use only

Bugs

  • Saving a webpage from archive.org/wayback machine is broken.
  • Certain sites might have broken java script

Legal Disclaimer

1. No Warranty

This software is provided “as is”, without warranty of any kind, express or implied, including but not limited to the warranties of merchantability, fitness for a particular purpose, and noninfringement. In no event shall the authors or copyright holders be liable for any claim, damages, or other liability arising from, out of, or in connection with the software or the use or other dealings in the software.

2. User Responsibility

Users are solely responsible for how they use this software. This includes, but is not limited to, ensuring compliance with all applicable laws, regulations, and third-party terms of service when archiving, storing, or distributing content.

3. Copyright and Content Ownership

This software may be used to archive publicly accessible web content. However, the ownership and copyright of archived materials remain with their respective owners. Users must ensure they have the legal right to archive and store such content and must respect copyright, licensing, and intellectual property laws.

4. No Endorsement or Affiliation

Archived content does not imply any endorsement, sponsorship, or affiliation with the original content creators, websites, or organizations. This project is an independent tool and is not associated with any third-party services or platforms.

5. Compliance with Website Policies

Users are responsible for respecting website-specific rules, including robots.txt directives, terms of service, and rate limits. The developers of this software are not responsible for misuse that violates such policies.

6. Data Storage and Security

Users are responsible for securing any archived data and for ensuring that sensitive or personal information is handled in accordance with applicable data protection and privacy laws.

7. Limitation of Liability

Under no circumstances shall the developers or contributors be liable for any direct, indirect, incidental, special, consequential, or exemplary damages resulting from the use or inability to use this software.

8. Changes to This Disclaimer

The maintainers reserve the right to modify this disclaimer at any time without prior notice.


By using this software, you acknowledge that you have read, understood, and agree to this disclaimer.

About

A docker container to crawl and archive an entire domain with an inbuilt archive viewer.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages