I'm a data scientist who builds small AI tools and decision systems. I like work where the reasoning stays visible: what the data says, what the system did, and where a person should still decide.
- Loop Agent — a small agent loop with safe tools, SQLite memory, and a step-by-step trace. It runs locally without an API key.
- Point2Prompt — point at a UI element and turn feedback into a structured brief for a coding agent.
- SplitTaste — a reproducible experiment for repairing mixed-household streaming recommendations.
- Where to Sit — a seat-view simulator for making a better cinema seat choice.
- Microsoft Agent Framework — exposed authorization URLs for Work IQ tool sources and added regression coverage.
- AO Bench — fixed missing or empty benchmark reports; the change was merged with a passing CI matrix.
- AO Bench — added a runnable comparison walkthrough for two agent adapters.
I'm interested in agent reliability, useful evaluation, and human-in-the-loop product design.