- Overview
- Problem Statement
- About SetFit
- Project Architecture
- Methodology
- Pipeline
- Experiments & Metrics
- How to Run
- Key Insights
- Results
- Future Work
- Authors
This project implements and extends SetFit (Sentence Transformer Fine-tuning) for few-shot text classification, along with advanced improvements such as hard negative mining and open-set recognition.
We reproduce the methodology proposed in L. Tunstall et al., "Efficient Few-Shot Learning Without Prompts", 2022 and build upon it with practical enhancements.
Traditional NLP models require large labeled datasets, which are:
- Expensive
- Time-consuming
- Domain-specific
π Our goal:
Build a high-performance text classifier using only a few labeled examples (e.g., 8 per class).
The SetFit paper proposes a prompt-free few-shot learning framework that avoids the limitations of prompt-based methods.
- Uses Sentence Transformers instead of large LLMs
- Performs contrastive learning using sentence pairs
- Trains a lightweight classification head on embeddings
- Requires no prompts or verbalizers
π The method:
- Fine-tune sentence embeddings using positive/negative pairs
- Train classifier on embeddings
- Comparable performance to large models
- Orders of magnitude faster training
- Works with very small datasets (few-shot)
- No prompt engineering needed
setfit-project/
β
βββ data/ # Saved dataset samples (.txt)
βββ results/ # Experiment outputs
β βββ baseline/
β βββ hard_negative/
β
βββ configs/
β βββ default.yaml # Centralized project settings
β
βββ src/
β βββ data_loader.py # Dataset + few-shot sampling
β βββ pair_builder.py # Pair construction
β βββ hard_negative.py # Hard negative mining
β βββ train.py # Training pipeline
β βββ evaluate.py # Metrics + plots
β βββ demo.py # Interactive inference
β βββ config_utils.py # YAML config loader
β
βββ requirements.txt # Project dependencies
βββ LICENSE # Apache license
βββ README.md # This file- Sample k examples per class
- Create positive + random negative pairs
- Train Sentence Transformer with CosineSimilarityLoss
- Train Logistic Regression classifier
We extend SetFit with:
Instead of random negatives:
- Select semantically similar but incorrect samples
- Forces model to learn fine-grained boundaries
π Improves robustness and generalization
- Balanced positive/negative sampling
- Hard + easy negatives mix
We introduce "OTHER" class using confidence thresholding
If model confidence < threshold:
β Reject prediction (unknown class)
We use:
- Business
- Entertainment
- Politics
- Sports
- Technology
π More realistic than AG News
Dataset β Few-shot Sampling β Pair Generation
β Sentence Transformer Fine-tuning
β Embedding Extraction
β Classifier Training
β Evaluation + Demo
We evaluate:
- Baseline SetFit
- Hard Negative SetFit
Across:
- Multiple seeds
- Few-shot setting (k=8)
- Using multiple classification metrics beyond accuracy for a more comprehensive evaluation.
We evaluate using standard classification metrics:
- Accuracy
- Precision (macro)
- Recall (macro)
- F1-score (macro)
- Mean Β± Standard Deviation across seeds
All defaults are now stored in:
configs/default.yamlYou can edit dataset/model/threshold/seeds and other values there.
All scripts support --config and still allow CLI overrides for key options.
pip install -r requirements.txtpython src/data_loader.py --config configs/default.yamlOutputs:
data/train.txt
data/test.txtpython src/train.py --config configs/default.yaml --mode baselinepython src/train.py --config configs/default.yaml --mode hard_negativepython src/evaluate.py --config configs/default.yaml --results_dir results/baseline --output_dir results/baseline
python src/evaluate.py --config configs/default.yaml --results_dir results/hard_negative --output_dir results/hard_negativepython src/demo.py --config configs/default.yaml --model_dir results/hard_negative/seed_0 --interactiveUse the dedicated research config and matrix runner:
python src/run_research_experiments.py --config configs/research.yaml --dry_run
python src/run_research_experiments.py --config configs/research.yamlRun a pilot subset before full execution:
python src/run_research_experiments.py --config configs/research.yaml --datasets bbc_news,ag_news --k_values 8 --pair_strategies random,hard --seeds 0,1,2Pair strategy options:
randomeasyhardmixed
Threshold analysis for uncertainty-aware behavior:
python src/threshold_analysis.py --config configs/research.yaml --model_dir results/research/bbc_news/k_8/hard/seed_0Input: "Stock markets crash globally"
Prediction: BUSINESS (0.82)
Input: "Aliens discovered on Mars"
Prediction: OTHER
-
Few-shot learning is highly unstable across seeds
-
Hard negatives improve:
- decision boundaries
- semantic understanding
-
Confidence thresholding enables real-world deployment
| Seed | Accuracy |
|---|---|
| 0 | 0.716 |
| 1 | 0.699 |
| 2 | 0.710 |
Mean Accuracy: ~0.708
| Seed | Accuracy |
|---|---|
| 0 | 0.955 |
| 1 | 0.935 |
| 42 | 0.939 |
Mean Accuracy: ~0.943
- Absolute gain: +23.5%
- Hard negative mining significantly improves:
- semantic discrimination
- decision boundary sharpness
π This demonstrates that pair quality > model complexity in few-shot learning.
- Adaptive hard negative mining
- Label semantic injection
- Contrastive loss variants
- Calibration & uncertainty estimation
- Mozeel Pradip Vanwani
- Mayank Seth
- Krishnkant Sahu
- Burri Vivek Vardhan Verma