This project implements a complete NLP Pipeline from scratch, focusing on subword tokenization and Named Entity Recognition (NER).
Modern LLMs rely on subword tokenization to handle rare words and cross-domain terminology. This repository demonstrates a custom implementation of Byte Pair Encoding (BPE), integrated into a Bi-LSTM neural network for sequence labeling tasks.
The system is built in two distinct modules:
A pure Python implementation of the BPE algorithm.
- Training: Iteratively learns merge rules from raw text to construct a vocabulary.
- Encoding: Tokenizes unseen text by applying learned merges (handling Out-Of-Vocabulary words effectively).
- Efficiency: Optimized for variable vocabulary sizes.
A comprehensive script containing the Deep Learning pipeline built with PyTorch.
- Model: Bidirectional LSTM (Bi-LSTM) to capture context from both directions.
- Dataset: Custom
NERDatasetclass that handles the alignment between BPE sub-tokens and word-level entity labels. - Embeddings: Learned embeddings for the custom BPE vocabulary.
Custom-BPE-Tokenizer/
├── src/
│ ├── base_tokenizer.py # Abstract Base Class
│ └── bpe_tokenizer.py # Core BPE Algorithm
├── data/ # Raw corpora and Tagged NER data
├── tokenizers/ # Directory for saved tokenizer artifacts
├── models/ # Directory for saved model checkpoints
├── train.py # Script to train the BPE Tokenizer
├── train_ner.py # Main script (Model, Dataset, Training Loop)
└── evaluate.py # Script to evaluate performance (F1 Score)
pip install -r requirements.txtTrain a BPE tokenizer on a raw text corpus.
python train.py --data data/raw/domain_1_train.txt --output tokenizers/my_bpe.pkl --vocab_size 5000Train the Bi-LSTM model using the custom tokenizer.
python train_ner.py --tokenizer tokenizers/my_bpe.pkl \
--train data/ner/train_1_binary.tagged \
--dev data/ner/dev_1_binary.tagged \
--output models/best_model.pt \
--epochs 5Run a full evaluation on the test set.
python evaluate.py --tokenizer tokenizers/my_bpe.pkl \
--test_data data/ner/dev_1_binary.tagged \
--model models/best_model.pt- Dynamic Padding: Uses
collate_fnto handle variable-length sequences efficiently in PyTorchDataLoader. - Robust Alignment: Solves the challenge of aligning word-level tags (e.g., "New York" -> LOC) with subword tokens (e.g., "New", "Y", "ork").
- Cross-Domain Ready: The modular design allows training tokenizers on specific domains to improve downstream performance.
Author: Itai Nulman
Data & Information Engineering Student @ Technion