CachedDupeScanner is an Android duplicate file scanner. It scans very large directories, persists metadata and hashes in a Room-backed cache, and accelerates subsequent scans by reusing unchanged results.
- Incremental scans: unchanged files are not re-hashed; cache is reused.
- Deferred hashing: SHA-256 is computed only when size collisions exist.
- Persistent cache: metadata stored in scan-cache.db (Room/SQLite), with stable numeric file identities, indexed 32-byte SHA-256 storage shared by derived data, and user-approved blocking upgrades before the app opens its data screens.
- Target management: save multiple scan targets; run per-target or batch scans.
- Duplicate grouping: database-backed result browsing with infinite scrolling for large datasets.
- Trash flow: move files to .CachedDupeScanner/trashbin with restore and permanent-delete controls. The bin is excluded from scans by default.
- Manage duplicates: group-detail views with multi-select, select-all, and specific delete tracking.
- Scan reports: timings, phase durations, hash candidate counts.
- App-wide task monitoring: in-app banners and a draggable bubble follow long-running work across screens, while Android notifications expose active work outside the app. Task surfaces show processing speed, elapsed time, and estimated remaining time.
- Background execution: an app-owned task runtime, foreground data-sync service, and partial WakeLocks keep supported work active across UI lifecycle changes and device sleep.
- Performance controls: separate configurable 1-32 worker limits for scan hashing and similarity feature calculation, optional memory usage overlay, shared RAM thumbnail retention, and configurable thumbnail/timeline preview sizing for heavy workloads.
- Rich media previews: Timeline video preview mode with a dedicated RAM cache policy, width snapping, and multi-line frame rows.
- Smart filters: Saved filters, per-rule any/all member matching, same-folder and same-size group rules, similarity same-resolution and average-duration rules, and modified-time rules that persist across sessions.
- Advanced bulk delete: Full-candidate previews support keeping exactly one, at least one, or an exact custom count of matching or non-matching files by text rule, or keeping the oldest, newest, shortest-duration, or longest-duration file, with modified-time fallback for equal video durations.
- DB maintenance: purge missing files, re-hash stale or missing entries, rebuild duplicate groups, and scope maintenance to all cached files, detected duplicate-result groups, or generated similarity groups. Actionable via notification-backed execution.
- Similarity: configure and browse named video/image similarity clustering generated from scan-cache data after scans; exact reduced thumbnails are grouped by indexed 32-byte SHA-256 values, while numeric file-ID joins, bounded parallel media feature extraction, persistent sort options, and progress-tracked Update/Rebuild actions keep large result sets practical.
For filesystem scans:
- Collect eligible files:
FileWalkergathers metadata while applying configured exclusions. - Select candidates: non-zero files become hash candidates when their size collides in the current scan or persisted cache.
- Check the cache:
CacheStoreclassifies candidate metadata as FRESH, STALE, or MISS. - Hash only when needed: uncached, stale, or missing-hash candidates are processed through the configured bounded SHA-256 worker pool.
- Persist results: file metadata is cached and duplicate groups are derived from matching hashes.
The cache is designed to delay hashing as long as possible.
- Record eligible metadata: path, size, and modified time are stored for files that pass scan exclusions and cache settings; zero-size cache entries are skipped by default.
- Hash collisions during scans: regular scans hash non-zero files only when their size collides with another current or cached entry.
- Detect changes and repair on demand: size or modified-time changes mark entries as stale, while explicit DB maintenance can repair stale or missing hashes.
Compose UI
-> App task coordinator
-> Foreground service / notifications / partial WakeLocks
-> Scan engine
-> File walker
-> Bounded SHA-256 workers
-> Similarity / DB maintenance / Trash operations
-> Room cache and derived result tables
Key design points:
- Deferred hashing to minimize CPU and I/O
- Bounded workers and chunked writes for large datasets
- Stable numeric file identities for derived-table joins, with normalized paths retained for unique lookup and display
Room database (scan-cache.db) core tables:
- cached_files: path, size, mtime, hash
- scan_reports: scan summary (durations, counts, targets)
- trash_entries: trash records (origin/trashed path, size, timestamps)
- dupe_groups: materialized snapshot of duplicate groups for fast paginated browsing
- similarity_settings / similarity_setting_files / method-specific feature tables / similarity_clusters / similarity_cluster_members: configured similarity methods, file state, compact method-specific feature storage including binary thumbnail hashes, and sidecar member links for similarity-based duplicate candidates
- core:
FileMetadata,ScanResult, duplicate analysis - engine:
IncrementalScanner,FileWalker, hashing - cache: Room entities/DAO, cache lookup/upsert
- storage: settings, targets, reports, Trash, and Similarity repositories
- tasks / notifications: app-owned task state, foreground execution, and Android notifications
- export: JSON/CSV result serializers used as a developer utility; they are not currently exposed in the app UI
- ui: Compose screens and components
- Dashboard (entry point)
- Permission (file access)
- Targets (scan targets)
- Scan Command (run scans + progress)
- Results / Files (duplicates and file list)
- Trash (restore/permanent delete)
- DB Management (cleanup/rehash)
- Similarity (managed similarity clustering and results)
- Reports (scan reports)
- Settings / About
- Filesystem scans can require broad storage access so the app can enumerate, hash, move, restore, and delete user-selected files.
- Scanning, hashing, cache maintenance, and Similarity processing run locally. The app does not declare the Android Internet permission.
- Metadata, hashes, reports, Trash records, and Similarity results are stored in the local
scan-cache.dbdatabase. - Normal deletion moves files to
.CachedDupeScanner/trashbin; permanent deletion from Trash cannot be undone. - Long-running work uses a foreground data-sync service and may display Android notifications.
- Android Studio or a command-line Android SDK installation with SDK Platform 36
- JDK 17 or newer to launch Gradle; the build selects a JetBrains JDK 21 daemon and Kotlin toolchain
- An Android API 24 or newer device or emulator for installation and instrumented tests
- Network access on the first build for Gradle dependencies and toolchain provisioning
./gradlew assembleDebug./gradlew testConnect an API 24 or newer device or start an emulator before running:
./gradlew connectedAndroidTestFor signed APK automation, see docs/android-apk-release.md.
Project rules and the agent guide are in AGENTS.md.
See LICENSE.