I build AI systems from the problem statement to production and beyond.
Getting the model to work is rarely the hard part. Defining the right problem is harder, and keeping the system trustworthy once real users depend on it is harder still.
Currently an AI engineer at Potloc in Montreal. Before that, Lead NLP Engineer at Valital, running the NLP roadmap for AML and KYC and setting up the evaluation and monitoring practice the team ran on.
-
decision-models-under-pressure An independent benchmark of seven decision models under three production stresses: growing candidate lists, shuffled option order, and harder distractors. Pre-registered before the first API call, with the dataset and per-call ledger published. Shuffling the answer options alone changed 14.6% of one model's decisions, and a fixed order does not fix it. The study also reversed the conclusion I started with, and the repo keeps the original reasoning rather than quietly rewriting it.
-
Assessing LLMs' Ability to Navigate Cultural Knowledge Conflicts (C3NLP 2024, non-archival). QARV, a 671-question benchmark for how LLMs handle conflicts between U.S. and Korean perspectives.
-
CLaC at SemEval-2020 Task 5 Multi-task stacked Bi-LSTMs for detecting the span of antecedents and consequents in counterfactual statements.
- MSc Computer Science, Concordia University. Thesis on input representations and classifiers for relation extraction, across SemEval-2010 Task 8, TACRED, Re-TACRED, and BioCreative VII (DrugProt).
- BEng Computer Engineering, Hongik University
English and Korean. Reachable on LinkedIn
