Independent Data & AI Engineer specializing in data analysis, advanced data cleaning, machine learning, and the development of artificial intelligence systems.
I build high-quality datasets, reproducible notebooks, predictive models, AI agents, and educational projects across multiple domains. My work primarily focuses on public data, institutional systems, socioeconomic trends, AI agents, and language models built from scratch.
My work includes:
- integrating, cleaning, and validating heterogeneous datasets;
- building comparable indicators and time series;
- developing transparent and reproducible analyses;
- auditing and monitoring data quality;
- building predictive models and decision-support systems;
- designing AI agents with routing, planning, tool use, and memory;
- progressively implementing language models and Transformers from scratch.
I primarily use Python to transform complex data into reliable, understandable, and actionable insights. I also build AI systems from the ground up to explore their underlying mechanisms in a transparent and practical way.
- Datasets Expert: ranked 77th out of 10,795 — highest rank: 70th — 13 medals (9 silver and 4 bronze)
- Notebooks Expert: ranked 535th out of 61,632 — highest rank: 515th — 18 bronze medals
| Project | Area | Tools and Key Features |
|---|---|---|
| Customer Support Agent | AI Agent / Human-in-the-Loop | Google ADK, Gemini, email classification, and human escalation |
| Building AI Agent | AI Agent from Scratch | Standard Python, routing, planning, parsing, tool execution, and memory |
| Building LLM | Language Model from Scratch | A progressive journey from a character-level statistical model to a small decoder-only Transformer |
| Global Inequality and Poverty — 1980–2024 | Socioeconomic Analysis | Data integration and development of globally comparable indicators |
| Italian Justice System Workload | Institutional Data | Civil and criminal justice analysis for 2003–2024, data auditing, and quality control |
| Home Credit Default Risk | Credit Risk / Machine Learning | XGBoost, LightGBM, SHAP, and feature engineering |
| Global Emissions & Temperature — 1950–2024 | Climate and Time Series | Analysis of CO₂ emissions, greenhouse gases, and global temperatures |
- Data quality and advanced data cleaning
- Applied machine learning
- Explainable AI
- AI agents and automation
- Natural Language Processing
- Language models and Transformers
- Artificial intelligence from scratch
- Public and institutional data
- Socioeconomic analysis
- Time series and comparable indicators
- Reproducible research