Skip to content
View felipeandrade91's full-sized avatar

Block or report felipeandrade91

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
felipeandrade91/README.md

Hi, I'm Felipe Andrade

Data Analyst | Data Scientist | SQL | Python | Power BI | Machine Learning | Customer Analytics

I'm a Data Analyst and Data Scientist with a PhD, focused on transforming complex datasets into analytical insights, predictive models, and data-driven solutions.

My background in scientific research has given me extensive experience in quantitative analysis, statistical reasoning, hypothesis-driven problem solving, and reproducible analytical workflows.

Today, I apply these skills to business problems across Data Analytics, Business Intelligence, Customer Analytics, Statistics, and Machine Learning.

I also have experience building scalable data pipelines and lakehouse architectures with Databricks, PySpark, and Delta Lake, allowing me to work across the path from raw data to analytical and machine learning applications.

Applying scientific rigor and analytical thinking to solve business problems through modern data analytics and data science.


Tech Stack

Data Analytics & Business Intelligence

  • SQL
  • PostgreSQL
  • Power BI
  • DAX
  • Power Query
  • Analytics Engineering
  • Data Modeling
  • Star Schema
  • Customer Analytics
  • Customer Segmentation
  • Customer Lifetime Value (CLV)
  • Data Visualization

Data Science & Machine Learning

  • Python
  • Pandas
  • Scikit-learn
  • XGBoost
  • Statistical Analysis
  • Statistical Inference
  • A/B Testing
  • Causal Inference
  • Predictive Modeling
  • Classification
  • Regression
  • Time Series Forecasting
  • Feature Engineering
  • Model Evaluation

Data Engineering & Deployment

  • PySpark
  • Databricks
  • Delta Lake
  • ETL / ELT
  • Data Quality
  • Medallion Architecture
  • Incremental Processing
  • FastAPI
  • REST APIs
  • Docker
  • Pytest

Tools

  • Git
  • GitHub
  • DBeaver
  • Excel

Featured Projects

⭐ Customer Analytics for Brazilian E-commerce

End-to-end Analytics Engineering and Business Intelligence project built with PostgreSQL, SQL, and Power BI.

The project implements a Star Schema data model, SQL analytical layer, business KPIs, data quality validation, and interactive dashboards covering sales, customers, logistics, and satisfaction.

🔗 Repository


⭐ Customer Churn Prediction

End-to-end Machine Learning and Customer Analytics project focused on predicting customer churn and translating model results into business insights.

The project covers SQL-based data preparation, exploratory data analysis, feature engineering, statistical analysis, classification modeling, model comparison, and business-oriented interpretation using Logistic Regression, Random Forest, and XGBoost.

The trained model was subsequently extended into a containerized REST API using FastAPI, Docker, Pydantic, and Pytest.

🔗 Repository 🔗 Deployment & API


⭐ Customer Segmentation & Customer Lifetime Value Analytics

Customer Analytics project built with PostgreSQL, SQL, and Python, focused on transforming transactional data into customer-level insights.

The project applies RFM segmentation, customer feature engineering, Historical Customer Lifetime Value (CLV), revenue concentration analysis, and business-oriented visualization.

🔗 Repository


⭐ Causal Inference & Experimentation

Data Science and statistical modeling project evaluating the incremental impact of digital advertising.

The project combines experimental and observational approaches using A/B Testing, Propensity Score Matching, Inverse Probability Weighting, Double Machine Learning, and Causal Forests to estimate treatment effects and investigate treatment heterogeneity.

🔗 Repository


⭐ Sales Forecasting with Machine Learning — Rossmann Stores

End-to-end Machine Learning project for retail sales forecasting.

The project covers exploratory time series analysis, temporal feature engineering, regression modeling, model comparison, and interpretation using Linear Regression, Random Forest, and XGBoost.

🔗 Repository


⭐ NYC Taxi Data Engineering Platform

Scalable data engineering project built with Databricks, PySpark, and Delta Lake, processing more than 38 million NYC Yellow Taxi trips.

The project implements a Medallion Architecture with Bronze, Silver, and Gold layers, including incremental processing, data quality validation, dimensional modeling, and analytical data preparation.

This project demonstrates my ability to work with the data infrastructure supporting large-scale analytics and machine learning workflows.

🔗 Repository


📂 Data Analytics & Data Science Portfolio

A curated collection of projects covering Data Analytics, Business Intelligence, Customer Analytics, Statistical Analysis, Machine Learning, Data Engineering, and Scientific Data Science.

The portfolio includes additional projects involving SQL, Python, Power BI, biodiversity data, exploratory analysis, statistical modeling, and data pipelines.

🔗 Explore the Portfolio


Scientific Background

My transition into business-oriented data work builds on a long-standing background in quantitative scientific research.

  • PhD in Animal Biology — UNICAMP
  • Postdoctoral Researcher — USP
  • 28 peer-reviewed scientific publications
  • Description of 13 new amphibian species
  • 15+ years working with complex real-world datasets
  • Extensive experience in statistical analysis, quantitative research, and reproducible workflows

This background provides a strong foundation in analytical thinking, statistical reasoning, scientific problem solving, and working with complex datasets.


Connect with Me

💼 LinkedIn

💻 GitHub

Pinned Loading

  1. Brazilian-Anuran-Diversity-Dashboard Brazilian-Anuran-Diversity-Dashboard Public

    Interactive Power BI dashboard exploring the spatial and temporal patterns of Brazilian anuran diversity using GBIF occurrence records.

    1

  2. gbif-amphibians-etl-pipeline gbif-amphibians-etl-pipeline Public

    This project was developed as a portfolio piece focused on biodiversity data engineering, demonstrating skills in SQL-based ETL pipelines, data quality assessment, and analytical dataset construction.

    1

  3. GBIF-Brazilian-Amphibian-Biodiversity-Analysis GBIF-Brazilian-Amphibian-Biodiversity-Analysis Public

    This project presents an exploratory and statistical analysis of amphibian occurrence records in Brazil using data from the Global Biodiversity Information Facility (GBIF).

    Jupyter Notebook 1

  4. Customer-Analytics-for-Brazilian-E-commerce Customer-Analytics-for-Brazilian-E-commerce Public

    End-to-end Customer Analytics project using PostgreSQL and Power BI, featuring Analytics Engineering, Star Schema modeling, SQL semantic layers, and interactive business dashboards built from the B…

    1