A multi-stage NLP pipeline designed to analyze environmental sentiment and extract representative consensus from public comments. The project leverages fine-tuned transformer models for classification and summarization, combined with semantic clustering for representative insight extraction.
File: sentiment_summarizer_app.py
The primary interactive tool. Input a list of raw comments and the AI will:
- Sentiment Analysis: Classify each comment as Positive or Negative using a fine-tuned RoBERTa model.
- Centroid Ranking: Identified the most representative "centroid" comments for each group using MiniLM embeddings.
- Consensus Summarization: Generate a concise summary of the group's collective opinion using a fine-tuned FLAN-T5 model.
To Run:
streamlit run sentiment_summarizer_app.pyFile: dashboard_app.py
A visualization suite for large-scale batch analysis (requires model_output.csv).
- Sentiment distribution and over-time trends.
- Interactive Word Clouds and composición analysis.
- Raw data explorer.
To Run:
streamlit run dashboard_app.pyModel_ROBERTA-Sentiment/: Fine-tuned model for environmental sentiment classification.Model_FLAN-T5/: Fine-tuned model for high-fidelity summarization.utils.py: Centralized logic for model loading and inference pipelines.datasets/: Training and testing data, including thesummarization_dataset.csvfor sample testing.tests/: Standalone CLI scripts for testing individual model performance (test_sentiment.py,test_summ.py).requirements.txt: Full dependency list.
-
Sentiment Classification: Based on
cardiffnlp/twitter-roberta-base-sentiment, fine-tuned on climate-specific datasets. -
Ranking (Centroid): Uses
all-MiniLM-L6-v2to vectorize comments. We calculate the mean vector (centroid) of each cluster and select the top$k$ comments with the highest cosine similarity to the centroid. -
Summarization: A sequence-to-sequence model (FLAN-T5) fine-tuned with specific prompts:
- Positive: "Summarize the opinions of users who believe in the reality of climate change.:"
- Negative: "Summarize the opinions of users who skeptical of/deny climate change.:"
- Create and activate a environment (Conda or venv):
conda create -n sentiment-analysis python=3.12
conda activate sentiment-analysis- Install dependencies:
pip install -r requirements.txt- Ensure models are placed in the root directory (refer to the Model Files section in structure).
The core models were developed and analyzed in the following notebooks:
Model_Training.ipynb: Original RoBERTa fine-tuning process.FLANT5-FT.ipynb: Dataset preparation and fine-tuning for the summarization model.Advanced_Sentiment_Analytics.ipynb: Initial exploration of keyword extraction and NER.
- Lawrence Agarin (lawrenceivanpagarin@iskolarngbayan.pup.edu.ph)
- Kyle Desmond Co (kyledesmondpco@iskolarngbayan.pup.edu.ph)
- Reymel Sardenia (reymelosardenia@iskolarngbayan.pup.edu.ph)
- Ken Satorre (kencalvinssatorre@iskolarngbayan.pup.edu.ph)
- Earl Andrei Fidel (earlandreidfidel@iskolarngbayan.pup.edu.ph)
🖥️ Bachelor of Science in Computer Science
Polytechnic University of the Philippines - Manila
Ria A. Sagum - Instructor (Natural Language Processing)
This project makes use of several open-source libraries and pre-trained models. We would like to acknowledge the contributions of the following:
- Hugging Face: For the
transformerslibrary and hosting the model hub. - Cardiff NLP: For the base
twitter-roberta-base-sentimentmodel. - Google Research: For the
FLAN-T5architecture. - Sentence Transformers: For the
all-MiniLM-L6-v2model used in semantic ranking. - Streamlit: For the interactive web application framework.
- Scientific Python Community: Including the developers of PyTorch, Pandas, NumPy, and Scikit-learn.
- Public Sentiment Analysis on Climate Change (simran98solanki): for the Twitter Sentiment Dataset.