I design evaluation benchmarks, LLM-judge rubrics, and adversarial test cases at Abundant (YC F24), finding exactly where agents fail before users do.
When I'm not breaking agents, I'm building them:
- 🤖 Real-time multimodal agent: vision + speech + prosody in one reasoning loop
- 🔍 RAG pipelines: on PostgreSQL with LangChain/LangGraph
- ⚡ Agent evals: judge calibration, failure-mode analysis, SWE-bench runs
🔭 Next on the bench: an MCP server + a finance-ops collections agent
💬 Ask me about: making LLM judges confess their biases
| Project | What it does | Stack |
|---|---|---|
| Multimodal Emotion Agent | Infers emotional/cognitive state from live webcam + mic in real time | PyTorch · YOLO · CLIP · Whisper · LLMs |
| SWE-bench Agent Evaluation | Ran SWE-agent on a Verified subset — 64% resolve rate + failure analysis | Python · SWE-agent · Git |
| RAG Trip-Planner Agent | Multilingual chatbot + trip planning over Postgres | LangChain · Gemini · PostgreSQL |
| Time-Series Forecasting Pipeline | Multivariate forecasting — benchmarked ~4 models, XGBoost champion / LSTM challenger, ~25% error reduction vs RNN baseline | XGBoost · LSTM · Python |
- AI/ML and GenAI projects
- Python backend or data-driven applications
- Research-oriented or open-source AI initiatives
I get paid to break AI agents. The agents are not amused. 🤖💥





