M.Sc. Computer Science · TU Darmstadt
B.Sc. Applied Computer Science and Artificial Intelligence · Sapienza University of Rome · Expected Dec 2026
Website · Hugging Face · LinkedIn · Email · Research CV
Primary focus: Information Retrieval, RAG/LLM Evaluation, and Reproducible ML Systems. Secondary direction: exact evaluation of reinforcement-learning agents.
I study information retrieval and retrieval-augmented generation, focusing on when improvements in retrieval do—or do not—lead to more accurate and grounded answers.
My work combines controlled retrieval and generation experiments with paired evaluation, observability, failure analysis, versioned manifests, and inspectable artifacts.
| Project | Focus | Current scope |
|---|---|---|
| rag-observatory · Live Space · Toy Dataset | Trace-based analysis for RAG systems | Research prototype for inspecting retrieved evidence, generated answers, execution traces, and failure labels |
| msmarco-genqa · Benchmark Runs | Retrieval-augmented generation on MS MARCO | Retrieval, reranking, generation, grounding analysis, paired statistical evaluation, and reproducible experiment reports |
| q-learning-exploitability · Evidence Map | Exact agent evaluation and controlled failure analysis | Adversarial backups and D4 evidence pooling reduced force-loss policies from 6/6 to 0/6 under matched Q-update budgets |
| Public research artifacts | Reusable evidence for evaluation work | Hugging Face datasets / Spaces, versioned reports, release archives, manifests, and trace examples connected back to the source repositories |
Only merged, publicly verifiable contributions are listed here.
| Ecosystem | Contribution | Evidence |
|---|---|---|
| MTEB | Fixed duplicate counting for symmetric STS pairs. | embeddings-benchmark/mteb#4958 |
| Pyserini | Fixed M-BEIR instruction lookup from cache-home paths. | castorini/pyserini#2655 |
| MTEB | Updated GermanGovService retrieval to the v2 dataset. | embeddings-benchmark/mteb#5323 |
| MTEB Leaderboard | Added model language-scope display on leaderboard cards. | embeddings-benchmark/leaderboard-frontend#31 |
Research question: When does better retrieval improve grounded generation, and when do conventional evaluation metrics hide the failure?
Current study: Cross-dataset retrieval and reranking evaluation on MS MARCO, TREC-DL, SciFact, and NFCorpus, including retrieval-depth sensitivity, candidate coverage, and query-level failure analysis.
Secondary direction: Exact evaluation of reinforcement-learning agents under worst-case opponents.




