Data Analyst | Data Scientist | SQL | Python | Power BI | Machine Learning | Customer Analytics
I'm a Data Analyst and Data Scientist with a PhD, focused on transforming complex datasets into analytical insights, predictive models, and data-driven solutions.
My background in scientific research has given me extensive experience in quantitative analysis, statistical reasoning, hypothesis-driven problem solving, and reproducible analytical workflows.
Today, I apply these skills to business problems across Data Analytics, Business Intelligence, Customer Analytics, Statistics, and Machine Learning.
I also have experience building scalable data pipelines and lakehouse architectures with Databricks, PySpark, and Delta Lake, allowing me to work across the path from raw data to analytical and machine learning applications.
Applying scientific rigor and analytical thinking to solve business problems through modern data analytics and data science.
- SQL
- PostgreSQL
- Power BI
- DAX
- Power Query
- Analytics Engineering
- Data Modeling
- Star Schema
- Customer Analytics
- Customer Segmentation
- Customer Lifetime Value (CLV)
- Data Visualization
- Python
- Pandas
- Scikit-learn
- XGBoost
- Statistical Analysis
- Statistical Inference
- A/B Testing
- Causal Inference
- Predictive Modeling
- Classification
- Regression
- Time Series Forecasting
- Feature Engineering
- Model Evaluation
- PySpark
- Databricks
- Delta Lake
- ETL / ELT
- Data Quality
- Medallion Architecture
- Incremental Processing
- FastAPI
- REST APIs
- Docker
- Pytest
- Git
- GitHub
- DBeaver
- Excel
End-to-end Analytics Engineering and Business Intelligence project built with PostgreSQL, SQL, and Power BI.
The project implements a Star Schema data model, SQL analytical layer, business KPIs, data quality validation, and interactive dashboards covering sales, customers, logistics, and satisfaction.
End-to-end Machine Learning and Customer Analytics project focused on predicting customer churn and translating model results into business insights.
The project covers SQL-based data preparation, exploratory data analysis, feature engineering, statistical analysis, classification modeling, model comparison, and business-oriented interpretation using Logistic Regression, Random Forest, and XGBoost.
The trained model was subsequently extended into a containerized REST API using FastAPI, Docker, Pydantic, and Pytest.
🔗 Repository 🔗 Deployment & API
Customer Analytics project built with PostgreSQL, SQL, and Python, focused on transforming transactional data into customer-level insights.
The project applies RFM segmentation, customer feature engineering, Historical Customer Lifetime Value (CLV), revenue concentration analysis, and business-oriented visualization.
Data Science and statistical modeling project evaluating the incremental impact of digital advertising.
The project combines experimental and observational approaches using A/B Testing, Propensity Score Matching, Inverse Probability Weighting, Double Machine Learning, and Causal Forests to estimate treatment effects and investigate treatment heterogeneity.
End-to-end Machine Learning project for retail sales forecasting.
The project covers exploratory time series analysis, temporal feature engineering, regression modeling, model comparison, and interpretation using Linear Regression, Random Forest, and XGBoost.
Scalable data engineering project built with Databricks, PySpark, and Delta Lake, processing more than 38 million NYC Yellow Taxi trips.
The project implements a Medallion Architecture with Bronze, Silver, and Gold layers, including incremental processing, data quality validation, dimensional modeling, and analytical data preparation.
This project demonstrates my ability to work with the data infrastructure supporting large-scale analytics and machine learning workflows.
A curated collection of projects covering Data Analytics, Business Intelligence, Customer Analytics, Statistical Analysis, Machine Learning, Data Engineering, and Scientific Data Science.
The portfolio includes additional projects involving SQL, Python, Power BI, biodiversity data, exploratory analysis, statistical modeling, and data pipelines.
My transition into business-oriented data work builds on a long-standing background in quantitative scientific research.
- PhD in Animal Biology — UNICAMP
- Postdoctoral Researcher — USP
- 28 peer-reviewed scientific publications
- Description of 13 new amphibian species
- 15+ years working with complex real-world datasets
- Extensive experience in statistical analysis, quantitative research, and reproducible workflows
This background provides a strong foundation in analytical thinking, statistical reasoning, scientific problem solving, and working with complex datasets.
💻 GitHub