Curriculum Vitae

Houman Rajabi

AI & ML Engineer

rajabi_houman@yahoo.com
+39 351 9465199
Milan, IT

Summary

AI & ML Engineer dedicated to translating business goals into scalable AI systems. My core focus is transforming inherently unpredictable AI models into reliable systems through disciplined evals, error analysis loops, and system architecture.I believe in a spec-first approach: whether shaping the build for a rapid MVP, orchestrating multi-agent systems, or steering AI coding assistants, I define evaluation criteria up front. Grounded in software engineering fundamentals, I optimize for architectural trade-offs (cost vs. latency) and production stability. I have worked in Italy, Germany, and Iran.

Work Experience

  • NLP Engineer (MSc thesis)
    2025-08 - 2026-01
    Badoo
    Developed privacy-first NLP pipelines and a novel pragmatic scoring metrics for unstructured data ingestion.
    • Engineered a custom metric that recovered 48.6% of valid user instances by fixing a critical failure mode where safety models over-penalized conversational text.
    • Built an end-to-end local LLM pipeline to synthesize contrastive examples, deploying a context-aware classifier immune to previous over-penalization flaws.
    • Fine-tuned BERT on user interaction graphs via contrastive triplet loss to extract latent intent and lifestyle features, feeding behaviorally-aligned embeddings into the core matching engine.
  • Research Scientist
    2023-09 - 2025-05
    DeepHealth Project
    Federated Learning and privacy-preserving AI for clinical and hospital environments.
    • Federated Learning: Architected decentralized training workflows across hospital networks to avoid moving raw patient data.
    • Technical Liaison: Aligned clinical requirements with HPC infrastructure for industrial partners like Philips and Thales.
    • Privacy Engineering: Deployed GDPR-compliant algorithms to ensure data sovereignty within strict hospital IT environments.
  • ML Engineer
    2021-03 - 2023-07
    Digikala Group
    Search & Recommendations
    • Engineered hybrid recommendation engines (Association Rules, Collaborative Filtering) to redesign the 'Frequently Bought Together' module, boosting Average Order Value (AOV) by 5% via cross-selling.
    • Developed price optimization algorithms for Flash Sales ('Shegeftane') to automate candidate selection and discount depth, achieving 95% sell-through without eroding margins.
    • Built robust demand forecasting pipelines (XGBoost, Prophet) capable of handling 4x traffic loads during Black Friday, reducing logistics bottlenecks by 10% and improving inventory accuracy.
    • Enhanced search relevance by integrating custom Persian NLP models for semantic matching and misspelling correction, reducing 'Zero Search Result' queries by 7%.
  • ML Engineer
    2019-05 - 2021-03
    Snapp! Marketplace Intelligence
    Marketplace intelligence and forecasting systems.
    • Forecasting: Developed models to balance driver supply and demand, reducing wait times by 2 minutes.
    • Algorithms: Designed dynamic pricing algorithms to maximize revenue (+3%) during high-demand periods.
    • Machine Learning: Utilized real-time data to enhance ETA prediction accuracy by 2.5%.
    • Analytics: Monitored KPIs to identify and execute optimizations that improved driver utilization and completion rates.
  • Project Contributor
    2018-09 - 2019-03
    Sharif University of Technology
    Institutional analytics and decision support dashboards.
    • Created analytical dashboards to support institutional decision-making and enhance strategic academic planning.

Education

  • Master of Science in Language Technologies (NLP)
    2023-09 - 2026-07
    University of Turin · MSc
    Focused on unstructured data and natural language processing, including semantic retrieval, multilingual NLP, and RAG systems.
    GPA: 110 / 110 cum laude | Weighted GPA: 28.8 / 30
    Courses: Semantic retrieval, Multilingual NLP, RAG systems, Unstructured data processing
  • Bachelor of Science in Computer Science
    2015-09 - 2019-03
    Sharif University of Technology · BSc
    Studied at one of the region's most competitive engineering institutions, with a strong emphasis on algorithmic logic and computational efficiency.
    Courses: Algorithms, Data structures, Software engineering, Computer science fundamentals

Projects

  • M.Sc. Thesis: The RLHF Modulation Paradox
    2026-02 - 2026-07
    Contrastive Synthetic Augmentation for LGBTQ+ Slur Reclamation
    University of Turin — Supervisors: Prof. Viviana Patti, Prof. Cristina Bosco, Prof. Livio Bioglio
    • Custom scoring module (MPEC): Fixed a failure mode where safety-aligned models over-weight surface toxicity tokens; the correction recovered 48.6% of valid instances the original metric silently discarded.
    • End-to-end pipeline: Five-stage local generation pipeline (Llama-3-8B), a zero-trust annotation platform with community-insider stratification, and a Chain-of-Thought validator ensemble — all designed and built from scratch.
    • Privacy-by-design architecture: All generation, scoring, and annotation ran fully on-premises — no community-register text was transmitted to third-party services. Annotators identified solely by cryptographic pseudonyms.
    Tools: Llama-3-8B, BERT, MPEC scoring, Label Studio
  • Open-Market-Intelligence
    RAG & Multimodal AI
    • Developed a multimodal RAG pipeline for financial analysis using Qwen2-VL for vision-based table extraction and vLLM for semantic chunking, incorporating sandboxed code execution to render accurate plots instead of hallucinating visuals.
    • Designed a retrieval and reasoning workflow for heterogeneous financial documents including PDFs, Excel reports, and embedded images.
    Tools: Qwen2-VL, vLLM, FAISS, sandboxed Python execution
  • Role-Sync
    Agentic AI & Orchestration
    • Built a human-in-the-loop resume analysis agent using LangGraph and Flask to map skills and identify gaps interactively.
    • Implemented conversational state management with approval flows to support real-world hiring workflows.
    Tools: LangGraph, Flask, OpenAI, structured state machines
  • EVALITA-MultiPRIDE-2026
    NLP Research & Fine-tuning
    • Designed a hybrid NLP framework (BERT + MLP) for detecting slur reclamation across three languages, utilizing language-specific feature fusion.
    • Built the evaluation pipeline used to validate the system, which ranked 1st out of 24 competing runs on the Italian subtask (Macro F1 = 0.8981).
    Tools: BERT, MLP, language-specific feature engineering, scikit-learn

Skills

Languages & Core

  • Python (OOP)
  • SQL
  • Bash
  • Git
  • Docker

MLOps & Cloud

  • AWS SageMaker
  • GCP Vertex AI
  • Azure ML
  • MLflow
  • CI/CD
  • HPC
  • Slurm

LLMs & GenAI

  • RAG pipelines
  • Hugging Face Transformers
  • AI Evals
  • LangChain
  • LangGraph
  • Vector Search
  • vLLM
  • Fine-tuning (PEFT/LoRA)
  • CUDA

Machine Learning

  • PyTorch
  • TensorFlow
  • NLP
  • Classification/Regression
  • Federated Learning
  • Recommendation Systems
  • Scikit-learn
  • XGBoost
  • Computer Vision

Data & Big Data

  • Vector DBs
  • PySpark
  • Apache Spark
  • Kafka
  • Databricks
  • NoSQL
  • Data Warehousing

Publications

  • Identity, Toxicity, or Complexity? A Language-Specific Feature Selection Approach to Reclamatory Intent Detection
    2026
    Proceedings of EVALITA 2026 (Task A: MultiPRIDE), Bari, Italy
    Published, First Author
    Rajabi, H., Ghavidel, F., & Ghahremani, K.
    1st Place — Italian subtask, 24 competing runs. Hybrid BERT + MLP with language-specific sociolinguistic feature fusion. Macro F1 = 0.8981.
  • Parametric Stubbornness: Mechanistically Isolating the Layer Shift and Sparsity Gradient of RAG Knowledge Conflicts in Llama-3
    2026
    Preprint. Zenodo
    Preprint
    Rajabi, H.
    Applied Two-Phase Activation Patching on Llama-3-8B across 452 minimal-pair RAG knowledge conflicts; demonstrated causally that hallucination is driven by contextual attention override rather than parametric suppression, and that implicit conflicts route through a sparse two-head semantic bottleneck.
  • Reclamation-Aware Augmentation: A Community-in-the-Loop Pipeline for Generating Contrastive Slur-Reclamation Data
    2026
    Manuscript in preparation
    Manuscript in preparation
    Rajabi, H.
    Five-stage multi-model pipeline producing contrastive minimal pairs for pragmatically sensitive classification; introduces MPEC, a masked composite scoring gate validated against community annotators (72%–100% confirmation monotonically by score band); 154 community-validated pairs double the reclamatory training class and raise downstream binary F1 by up to 17.6%.

Awards

  • 1st Place — EVALITA 2026
    2026
    Italian subtask, 24 competing runs
    Ranked #1 out of 24 competing teams on Italian reclamatory intent detection with Macro F1 = 0.8981.
  • Novel Scoring Metric (MPEC)
    2026
    Master's thesis
    Built a scoring component that recovered 48.6% of valid reclaimed pairs an existing metric missed.

Certificates

  • AWS Certified Machine Learning — Specialty
    AWS
    Hands-on end-to-end ML pipeline design, model deployment, and real-time inference on AWS.

Languages

  • English
    Full professional proficiency
  • Persian
    Native
  • German
    Limited working proficiency
  • Italian
    Elementary proficiency