A lightweight, self-hosted web archiving system. Archive any URL, subdomain, or entire domain — snapshots stored locally, viewable offline.
- Single URL, multi-URL, subdomain, or full-domain crawling
- One snapshot per URL — archiving again overwrites the previous
- Content-addressed asset storage — identical assets stored once via SHA-256 hash
- Playwright-powered crawling — handles JS-rendered pages
- Snapshot viewer — view archived pages in-browser with local assets
- Real-time job monitoring — live progress updates for active crawls
- SQLite database — zero-dependency metadata storage
- Fully Dockerized — one command to run
- Docker & Docker Compose
git clone https://github.com/nr-yolo/Web-Archiver.git
cd webarchive
mkdir data
docker compose up -d --buildOpen http://localhost:8777 in your browser.
┌─────────────────┐ ┌──────────────────────┐
│ React Frontend │────▶│ FastAPI Backend │
│ (Nginx :8080) │ │ (Uvicorn :8000) │
└─────────────────┘ │ │
│ ┌─────────────────┐ │
│ │ Playwright │ │
│ │ Crawler Engine │ │
│ └────────┬────────┘ │
│ │ │
│ ┌────────▼────────┐ │
│ │ SQLite Database │ │
│ │ /data/webarchive│ │
│ │ .db │ │
│ └────────┬────────┘ │
│ │ │
│ ┌────────▼────────┐ │
│ │ File Storage │ │
│ │ /data/snapshots │ │
│ └─────────────────┘ │
└──────────────────────┘
All endpoints are under /api/. No authentication required.
Start an archiving job.
{
"urls": ["https://example.com"],
"scope": "single",
"max_depth": 3,
"max_pages": 100
}Scopes:
single— archive only the submitted URLs, no crawlingsubdomain— crawl all pages within the same subdomaindomain— crawl the entire domain
Response:
{
"job_id": "uuid",
"status": "pending",
"urls": ["https://example.com"]
}List all archived snapshots.
Query params: q (search), limit, offset
View a snapshot. {url} should be URL-encoded.
Serve a stored asset (CSS, JS, image, font, etc.)
Delete a snapshot and its files.
List recent crawl jobs.
Get status of a specific job.
System statistics (snapshot count, asset count, storage size).
/data/
├── webarchive.db # SQLite metadata
└── snapshots/
├── _assets/ # Content-addressed assets
│ ├── ab/
│ │ └── ab3f9c....css
│ └── ...
├── https___example_com/
│ └── index.html # Archived page (assets rewritten)
└── https___another_com_page/
└── index.html
- Single: fetches only the given URL(s), no link following
- Subdomain: follows links within the same host (e.g.,
blog.example.com) - Domain: follows links within the same registered domain (e.g.,
*.example.com)
External CSS, JS, fonts, and images are fetched and stored locally (needed to render pages correctly). External HTML pages are not followed.
All assets are stored by SHA-256 hash. If two pages use the same CSS file, it's stored once. The HTML is rewritten to point to /api/asset/{hash}.
Links between archived pages are rewritten to /snapshot/{url} so navigation stays within the archive.
Environment variables (set in docker-compose.yml):
| Variable | Default | Description |
|---|---|---|
DB_PATH |
/data/webarchive.db |
SQLite database location |
STORAGE_PATH |
/data/snapshots |
Snapshot storage root |
cd backend
pip install -r requirements.txt
playwright install chromium
uvicorn main:app --reload --port 8000cd frontend
npm install
REACT_APP_API_URL=http://localhost:8000/api npm start| Layer | Technology |
|---|---|
| Frontend | React 18, React Router, react-hot-toast |
| Backend | FastAPI, Uvicorn |
| Crawler | Playwright (Chromium) |
| HTTP | aiohttp, httpx |
| HTML parsing | BeautifulSoup4, lxml |
| Database | SQLite via aiosqlite |
| Storage | Local filesystem |
| Serving | Nginx (frontend), Uvicorn (backend) |
| Container | Docker, Docker Compose |
- Crawling large sites can be slow — Playwright launches a full Chromium instance per crawl job
- For very large sites (1000+ pages), consider increasing
max_pagescarefully - The archive viewer opens snapshots in a new tab; internal links navigate between archived pages
- No authentication — intended for local/trusted network use only
- Saving a webpage from archive.org/wayback machine is broken.
- Certain sites might have broken java script
This software is provided “as is”, without warranty of any kind, express or implied, including but not limited to the warranties of merchantability, fitness for a particular purpose, and noninfringement. In no event shall the authors or copyright holders be liable for any claim, damages, or other liability arising from, out of, or in connection with the software or the use or other dealings in the software.
Users are solely responsible for how they use this software. This includes, but is not limited to, ensuring compliance with all applicable laws, regulations, and third-party terms of service when archiving, storing, or distributing content.
This software may be used to archive publicly accessible web content. However, the ownership and copyright of archived materials remain with their respective owners. Users must ensure they have the legal right to archive and store such content and must respect copyright, licensing, and intellectual property laws.
Archived content does not imply any endorsement, sponsorship, or affiliation with the original content creators, websites, or organizations. This project is an independent tool and is not associated with any third-party services or platforms.
Users are responsible for respecting website-specific rules, including robots.txt directives, terms of service, and rate limits. The developers of this software are not responsible for misuse that violates such policies.
Users are responsible for securing any archived data and for ensuring that sensitive or personal information is handled in accordance with applicable data protection and privacy laws.
Under no circumstances shall the developers or contributors be liable for any direct, indirect, incidental, special, consequential, or exemplary damages resulting from the use or inability to use this software.
The maintainers reserve the right to modify this disclaimer at any time without prior notice.
By using this software, you acknowledge that you have read, understood, and agree to this disclaimer.