A blazing-fast, multithreaded recursive directory traversal and file counting tool written in C with a Python interface. It is designed to scan large filesystems (millions of files across NVMe, SSD, HDD, APFS, NFS, or Lustre storage) with minimal overhead by interfacing directly with low-level kernel system calls and native multi-core worker threads.
Calling it just fast is an understatement. Here is an example of it counting almost 10Mil files in just over 30s.
(.venv) prasannakoirala@Prasannas-MacBook-Pro-258 fast_file_count % python fastcount.py / -t 8 -p
[Progress] Counted 1662063 entries (1339003 files, 323060 folders) [Queue: 8420] [332412 entries/sec]
[Progress] Counted 3102507 entries (2558820 files, 543687 folders) [Queue: 12510] [288088 entries/sec]
[Progress] Counted 4546325 entries (3827583 files, 718742 folders) [Queue: 15320] [288763 entries/sec]
[Progress] Counted 5983603 entries (5106868 files, 876735 folders) [Queue: 11200] [287455 entries/sec]
[Progress] Counted 7250787 entries (6231526 files, 1019261 folders) [Queue: 6340] [253436 entries/sec]
[Progress] Counted 8375269 entries (7230132 files, 1145137 folders) [Queue: 2410] [224896 entries/sec]
[Progress] Counted 9418223 entries (8039887 files, 1378336 folders) [Queue: 310] [208590 entries/sec]
Final Count:
Files: 8,284,344
Folders: 1,428,772
Total: 9,713,116 entries in 36.356 seconds (267,165 entries/sec)
Supports Linux (via direct SYS_getdents64 bulk kernel reads) and macOS / BSD (via multi-threaded d_type directory streams).
Doesn't support windows yet because I can't test it on windows.
- Features
- How It Works
- Prerequisites & System Requirements
- Step-by-Step Installation & Setup
- Usage Guide
- Python API Reference
- Performance Tuning & Best Practices
- Troubleshooting & FAQs
- Extreme Throughput: Traverses 300,000+ entries per second by avoiding userspace buffering and unnecessary
statcalls. - Cross-Platform: High-performance direct
SYS_getdents64on Linux with fully optimized multi-threaded POSIX fallback for macOS (APFS/HFS+) and BSD. - True Multi-Core Parallelism: Releases Python's GIL and utilizes native POSIX threads (
pthreads) with a scalable dynamic ring-buffer work queue. - Lock-Free Counters: Uses C11 atomic primitives (
stdatomic.h) to eliminate mutex contention during counting. - Ownership Filtering: Filter files by username or numeric UID on the fly.
- Live Progress Reporting: Optional background thread reports real-time count and traversal velocity (
entries/sec) every 5 seconds. - Resilient & Safe: Automatically elevates file descriptor limits and gracefully skips unreadable/permission-denied folders without crashing.
Traditional file counters (such as find . | wc -l, rsync, or Python's os.walk / os.scandir) introduce significant overhead from libc abstraction layers, single-threaded traversal, and per-file lstat() system calls. fastcount eliminates these bottlenecks:
[ Root Directory ]
│
(Work Queue) ◄────────────── Subdirectories Discovered
│
┌─────┴────────────────────────┐
│ Thread Pool (Workers) │
│ ┌─────────┐ ┌─────────┐ │
│ │Worker 1 │...│Worker N │ │
│ └────┬────┘ └────┬────┘ │
└───────┼─────────────┼────────┘
│ Bulk Dirent Buffer Read
▼
[ Parse directory entries ]
│
d_type == DT_DIR?
├── YES ──► Enqueue child path to Work Queue
└── NO / ANY ──► atomic_fetch_add(&g_total_files, 1)
- On Linux: Uses
syscall(SYS_getdents64, dir_fd, buf, 128KB)directly. This reads hundreds oflinux_dirent64directory records per kernel call in a single context switch. - On macOS / BSD: Uses per-thread directory streams with
d_typedirectly populated by APFS/HFS+, avoiding extralstatcalls.
Filesystems (ext4, XFS, Btrfs, APFS, ZFS) store the file type (DT_DIR, DT_REG, DT_LNK) directly in directory entries:
- If
d_type == DT_DIR, the subdirectory is immediately scheduled for traversal. fstatat()is only invoked as a fallback ifd_type == DT_UNKNOWNor when ownership filtering (--owner/UID) is requested.
- Employs a thread-safe circular ring buffer (
WorkQueue) synchronized with POSIX mutexes (pthread_mutex_t) and condition variables (pthread_cond_t). - Automatically doubles capacity if the queue becomes full.
- Whenever a worker encounters subdirectories, it enqueues them; idle workers wake up and dequeue directories independently.
- Automatically and cleanly terminates when the queue is empty and all active workers have finished.
File totals are recorded using C11 atomic operations (atomic_fetch_add(&g_total_files, 1)). Threads do not acquire locks to update counts, ensuring zero contention.
During directory traversal, the C extension calls Py_BEGIN_ALLOW_THREADS and Py_END_ALLOW_THREADS. This releases Python's Global Interpreter Lock (GIL), allowing all POSIX worker threads to run concurrently across all CPU cores.
On invocation, fastcount inspects the process file descriptor limit (RLIMIT_NOFILE) using getrlimit and automatically elevates rlim_cur to the maximum allowed limit (rlim_max or OS maximum), preventing descriptor exhaustion during deep scans.
- Operating System: Linux (x86_64, aarch64) or macOS (Apple Silicon M1/M2/M3/M4 or Intel).
- C Compiler: GCC (4.9+) or Apple Clang with C11 support (
stdatomic.h) andpthread. - Python: Python 3.6 or higher with C development headers.
- Build Tools:
setuptools,pip, andwheel.
# Ensure Xcode Command Line Tools are installed:
xcode-select --install
# Install Python (if not already installed):
brew install pythonsudo apt-get update
sudo apt-get install -y build-essential python3 python3-dev python3-pip python3-setuptoolssudo dnf groupinstall -y "Development Tools"
sudo dnf install -y python3 python3-devel python3-pip python3-setuptoolscd fast_file_count
ls -la
# Expected files: fastcount.c, fastcount.py, setup.pyChoose one of the following methods to build the C extension:
# Create and activate virtual environment
python3 -m venv .venv
source .venv/bin/activate
# Install build dependencies
pip install --upgrade pip setuptools
# Build the C extension in-place
python setup.py build_ext --inplacepip install .pip install -e .The repository includes a Python CLI wrapper fastcount.py.
usage: fastcount.py [-h] [-t THREADS] [-o OWNER] [-p] [-l] [-O OUTPUT] [-file | -all | -dir] path
Blazing fast recursive file counter and finder using parallel getdents64 / multithreaded dirent.
positional arguments:
path Target directory folder to traverse
options:
-h, --help Show this help message and exit
-t THREADS, --threads THREADS
Number of worker threads (default: CPU core count)
-o OWNER, --owner OWNER
Optional user ownership filter (username or numeric UID)
-p, --progress Print live progress every 5 seconds
-l, --list Dump discovered paths to stdout (like find)
-O OUTPUT, --output OUTPUT
Save dumped paths directly to specified file
-file, --file Dump files only when listing (default)
-all, --all Dump all entries (files + folders) when listing
-dir, --dir Dump folders/directories only when listing
Count all files and directories in a directory:
python3 fastcount.py /var/logScan large datasets or fast NVMe / multi-disk arrays:
python3 fastcount.py /data/dataset -t 16 -pOutput Example:
[Progress] Counted 2,289,656 entries (1,969,420 files, 320,236 folders) [Queue: 3,420] [457,931 entries/sec]
[Progress] Counted 2,778,174 entries (2,428,514 files, 349,660 folders) [Queue: 120] [97,703 entries/sec]
Final Count:
Files: 2,428,554
Folders: 349,660
Total: 2,778,214 entries in 10.009 seconds (277,571 entries/sec)
Stream all file paths to stdout for processing or piping:
# Stream files only (like find -type f)
python3 fastcount.py /data -l > filelist.txt
# Or explicitly pass -file
python3 fastcount.py /data -l -file > filelist.txt
# Pipe directly to tools like grep or wc
python3 fastcount.py /data -l | grep "\.parquet$"# Dump everything (files + folders)
python3 fastcount.py /data -l -all > all_entries.txt
# Dump folders only (like find -type d)
python3 fastcount.py /data -l -dir > folders.txtDirectly stream discovered paths into an output file with buffered writes:
# Dump files only to a file
python3 fastcount.py /data -O /tmp/files.txt -p
# Dump all entries (files + folders) to a file
python3 fastcount.py /data -O /tmp/all_entries.txt -all -pCount and list only files owned by a specific user:
python3 fastcount.py /srv/storage -o www-data -O www_files.txtWhen scanning system-wide mounts with restricted directories, run with sudo:
sudo python3 fastcount.py / -t 16 -pYou can import and use fastcount directly in your Python applications:
import fastcount
# Basic count
stats = fastcount.count_files("/path/to/directory")
print(f"Files: {stats['files']:,}")
print(f"Folders: {stats['dirs']:,}")
print(f"Total: {stats['total']:,}")
# Dump only files to a file
stats = fastcount.count_files(
path="/data/datasets",
threads=16,
uid=1000,
output_file="/tmp/dataset_files.txt",
type="f" # "f" for files (default), "d" for folders, "all" for both
)fastcount.count_files() returns a Python dictionary:
{
"files": int, # Total non-directory entries (regular files, symlinks, sockets, etc.)
"dirs": int, # Total subdirectories
"total": int # Grand total (files + dirs)
}| Parameter | Type | Default | Description |
|---|---|---|---|
path |
str |
Required | Absolute or relative root directory path to scan. |
threads |
int |
16 |
Number of worker threads spawned in the thread pool. |
uid |
int |
-1 |
Filter by user ID. Pass -1 (or omit) to disable filtering. |
progress |
bool |
False |
When True, a background thread logs progress to stdout/stderr every 5 seconds. |
list |
bool |
False |
When True, streams discovered paths to stdout (fd 1). |
output_file |
str |
None |
Path to save discovered paths to disk with 64KB buffered writes. |
type |
str |
"f" |
Entry filter: "f" (files only, default), "d" (folders only), "all", or "tagged". |
A high-speed directory comparison script built on top of fastcount. It scans source and destination directory trees in parallel, performs set-difference analysis on relative paths, prints summary statistics, and outputs a comprehensive list of missing and extra items.
- Parallel Scanning: Scans both
SRCandDESTtrees using multi-threadedfastcountkernel calls. - Categorized Breakdown: Tracks differences across files, folders, and grand totals.
- Comprehensive Diffs: Lists items in SRC missing from DEST, and extra items in DEST not in SRC.
- Export Lists: Optional
--save-diffsflag to savemissing_from_dest.txt,extra_in_dest.txt,in_both.txt, and summary reports.
usage: compare_files.py [-h] [-t THREADS] [-o OWNER] [-p] [-file | -all | -dir]
[-i IGNORE] [--src-ignore SRC_IGNORE] [--dst-ignore DST_IGNORE]
[-s SAVE_DIFFS] [-m MAX_DISPLAY]
src dest
positional arguments:
src Source directory path (SRC)
dest Destination directory path (DEST)
options:
-h, --help Show this help message and exit
-t THREADS, --threads THREADS
Number of worker threads (default: CPU core count)
-o OWNER, --owner OWNER
Optional user ownership filter (username or UID)
-p, --progress Print live progress every 5 seconds during scan
-file, --file Compare files only
-all, --all Compare both files and folders (default)
-dir, --dir Compare folders/directories only
-i, --ignore PATTERNS Comma-separated wildcard patterns ignored in both SRC and DEST
--src-ignore PATTERNS Comma-separated wildcard patterns ignored only in SRC
--dst-ignore PATTERNS Comma-separated wildcard patterns ignored only in DEST
-s, --save-diffs DIR Save full difference lists to text files in specified directory
-m, --max-display N Max difference lines to print to terminal (default: 100, 0 for all)
python3 compare_files.py /mnt/data_source /mnt/data_backup# Ignore temporary files, specific paths, and folders matching wildcards:
python3 compare_files.py /mnt/src /mnt/dest \
--src-ignore "/path/*/this,*.tmp,*cache*" \
--dst-ignore "*/backup/*,.DS_Store" \
-p
# Apply global ignore to both SRC and DEST:
python3 compare_files.py /mnt/src /mnt/dest -i "*.log,*.tmp,.git,node_modules"python3 compare_files.py /mnt/src /mnt/dest --save-diffs ./diff_reports -pExample Output:
[1/3] Scanning Source (SRC): /mnt/src ...
Ignore filters: *.tmp, */cache/*
Found 125,400 items (110,000 files, 15,400 folders, 420 ignored) in 0.220s
[2/3] Scanning Destination (DEST): /mnt/dest ...
Ignore filters: *.tmp, .git
Found 125,350 items (109,970 files, 15,380 folders, 1,200 ignored) in 0.232s
[3/3] Comparing relative directory structures ...
================================================================================
DIRECTORY COMPARISON SUMMARY
================================================================================
Source (SRC): /mnt/src
- Ignored: *.tmp, */cache/* (420 items skipped)
Destination (DEST): /mnt/dest
- Ignored: *.tmp, .git (1,200 items skipped)
Elapsed Time: 0.455s (SRC: 0.220s, DEST: 0.232s, Compare: 0.003s)
--------------------------------------------------------------------------------
Category Total Files Folders
--------------------------------------------------------------------------------
Total in SRC 125,400 110,000 15,400
Total in DEST 125,350 109,970 15,380
Present in BOTH (Matching) 125,300 109,950 15,350
In SRC, NOT in DEST (Missing) 100 50 50
In DEST, NOT in SRC (Extra) 50 20 30
================================================================================
[Saved Lists to './diff_reports']
- Missing from DEST: ./diff_reports/missing_from_dest.txt (100 lines)
- Extra in DEST: ./diff_reports/extra_in_dest.txt (50 lines)
- Present in BOTH: ./diff_reports/in_both.txt (125,300 lines)
- Summary Report: ./diff_reports/summary.txt
COMPREHENSIVE DIFFERENCE LIST
================================================================================
[+] In SRC but MISSING from DEST (100 items):
[FOLDER] logs/2026-08/
[FILE ] documents/report_final.pdf
[FILE ] images/cat1.png
[-] In DEST but NOT in SRC (Extra in DEST) (50 items):
[FOLDER] temp/old_cache/
[FILE ] temp/cache.tmp
- Thread Count (
-t):- Local NVMe SSDs / Apple Silicon APFS:
16to32threads usually saturate IOPS and maximize directory traversal speed. - Spinning HDDs:
4to8threads are recommended to prevent excessive disk head thrashing. - Distributed / Network Filesystems (NFS / Lustre / Ceph / GPFS): High thread counts (
32to64+) help hide network metadata round-trip latencies.
- Local NVMe SSDs / Apple Silicon APFS:
- Page Cache Warmth:
- Cold Cache: First run reads directory blocks from physical disk (limited by disk IOPS).
- Warm Cache: Subsequent runs traverse cached inode/dentry tables in RAM, scanning millions of entries in seconds.
- Permissions:
- Unreadable directories (e.g.
EACCES/EPERM) are skipped gracefully. To get a complete count across all system users, execute withsudo/rootpermissions.
- Unreadable directories (e.g.
A: Python development headers are missing. On Linux, install python3-dev (Ubuntu/Debian) or python3-devel (RHEL/CentOS/Rocky). On macOS, ensure Xcode command line tools are installed (xcode-select --install).
A: The username provided to -o / --owner could not be resolved in the system user database. Verify the username with id <username> or pass the numeric UID directly.
A: By design, symlinks to directories are not recursively followed to avoid infinite loops and double-counting across symlinked directory trees.