Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🔡 Custom BPE Tokenizer & NER System

Python PyTorch NLP

📋 Executive Summary

This project implements a complete NLP Pipeline from scratch, focusing on subword tokenization and Named Entity Recognition (NER).

Modern LLMs rely on subword tokenization to handle rare words and cross-domain terminology. This repository demonstrates a custom implementation of Byte Pair Encoding (BPE), integrated into a Bi-LSTM neural network for sequence labeling tasks.

🏗️ Architecture

The system is built in two distinct modules:

1. The Tokenizer (src/bpe_tokenizer.py)

A pure Python implementation of the BPE algorithm.

  • Training: Iteratively learns merge rules from raw text to construct a vocabulary.
  • Encoding: Tokenizes unseen text by applying learned merges (handling Out-Of-Vocabulary words effectively).
  • Efficiency: Optimized for variable vocabulary sizes.

2. The NER System (train_ner.py)

A comprehensive script containing the Deep Learning pipeline built with PyTorch.

  • Model: Bidirectional LSTM (Bi-LSTM) to capture context from both directions.
  • Dataset: Custom NERDataset class that handles the alignment between BPE sub-tokens and word-level entity labels.
  • Embeddings: Learned embeddings for the custom BPE vocabulary.

📂 Project Structure

Custom-BPE-Tokenizer/
├── src/
│   ├── base_tokenizer.py   # Abstract Base Class
│   └── bpe_tokenizer.py    # Core BPE Algorithm
├── data/                   # Raw corpora and Tagged NER data
├── tokenizers/             # Directory for saved tokenizer artifacts
├── models/                 # Directory for saved model checkpoints
├── train.py                # Script to train the BPE Tokenizer
├── train_ner.py            # Main script (Model, Dataset, Training Loop)
└── evaluate.py             # Script to evaluate performance (F1 Score)

🚀 How to Run

1. Setup Environment

pip install -r requirements.txt

2. Train the Tokenizer

Train a BPE tokenizer on a raw text corpus.

python train.py --data data/raw/domain_1_train.txt --output tokenizers/my_bpe.pkl --vocab_size 5000

3. Train the NER Model

Train the Bi-LSTM model using the custom tokenizer.

python train_ner.py --tokenizer tokenizers/my_bpe.pkl \
                    --train data/ner/train_1_binary.tagged \
                    --dev data/ner/dev_1_binary.tagged \
                    --output models/best_model.pt \
                    --epochs 5

4. Evaluate

Run a full evaluation on the test set.

python evaluate.py --tokenizer tokenizers/my_bpe.pkl \
                   --test_data data/ner/dev_1_binary.tagged \
                   --model models/best_model.pt

📊 Technical Highlights

  • Dynamic Padding: Uses collate_fn to handle variable-length sequences efficiently in PyTorch DataLoader.
  • Robust Alignment: Solves the challenge of aligning word-level tags (e.g., "New York" -> LOC) with subword tokens (e.g., "New", "Y", "ork").
  • Cross-Domain Ready: The modular design allows training tokenizers on specific domains to improve downstream performance.

Author: Itai Nulman Data & Information Engineering Student @ Technion

About

This project implements a complete NLP Pipeline from scratch, focusing on subword tokenization and Named Entity Recognition (NER)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages