Β Β Full stack software engineer with 10+ years of experience shipping production systems, now centered on AI-powered products: retrieval, LLM applications, recommendation systems, and agentic workflows.
Β Β I own products end-to-end β UI, APIs, data pipelines, model serving, cloud infrastructure, observability, and AI evaluation. Work has spanned systems serving millions of interactions, with 30%+ latency reductions and 2Γ+ inference throughput gains through model routing, batching, caching, and autoscaling.
Β Β I care about the unglamorous parts that make AI usable in production: groundedness and evaluation, cost per request, tracing and observability, and guardrails against prompt injection and tool abuse.
| Area | What I build |
|---|---|
| π€ Agentic AI | Multi-step agent loops, tool calling, structured outputs, MCP servers, human-in-the-loop approvals, long-running orchestration |
| π RAG & Retrieval | Hybrid search, embeddings, Graph-RAG, reranking, permissions-aware retrieval, long-context pipelines |
| β‘ LLM Serving & Cost | Unified model gateways, dynamic routing, provider failover, vLLM/Ray Serve, batching, caching, warm model pools |
| π Search & Recommendations | Learning-to-rank, personalization models, behavioral signals, real-time inference, A/B experimentation |
| π Data & ML Platform | Streaming/feature pipelines, embedding + indexing freshness, MLOps, deployment and rollout automation |
| π§ͺ Evaluation & Observability | Automated eval suites, hallucination/groundedness/tool-accuracy metrics, OpenTelemetry tracing, token & cost analytics |
| π AI Security | Prompt-injection defense, untrusted-content isolation, least-privilege tool access, RBAC/IAM, secrets management |
| π Full Stack Product | Streaming UIs, SSR/ISR apps, REST/GraphQL/gRPC APIs, event-driven microservices, payments and webhooks |
Languages
AI / ML & GenAI Engineering
RAG Graph-RAG Embeddings & Vector Search Agentic Workflows Multi-Agent Systems Agent Loops & Harnesses
Tool Calling Structured Outputs Context Engineering Prompt Engineering Multimodal AI Synthetic Data Generation
Model Routing LLM Evaluation Fine-Tuning (LoRA/QLoRA) MLOps ML Pipelines Pydantic LangSmith
Backend & APIs
REST WebSockets Server-Sent Events Microservices Serverless Event-Driven Architecture Distributed Systems
Concurrency Caching System Design DDD SOLID Design Patterns
Frontend & Full Stack
App Router SSR/ISR Zustand Streaming UI Responsive Design Accessibility Webhooks
Data, Search & Streaming
DynamoDB BigQuery ETL/ELT Streaming Pipelines Query Optimization Schema Design Feature Pipelines
Cloud & DevOps
EC2 ECS Lambda S3 RDS Bedrock SageMaker Helm Argo CD Ansible Infrastructure as Code CI/CD
Testing, Reliability & Observability
TDD Unit / Integration / E2E API Testing Coverage Gates Load & Performance Testing
Automated Evaluation Pipelines Distributed Tracing SLOs/SLIs Incident Response
AI Security & Application Security
OWASP Top 10 RBAC IAM OAuth 2.0 SSO/SAML Secrets Management Least-Privilege Tool Access
Prompt-Injection Defense Tool-Abuse Guardrails Untrusted-Content Isolation Secure Agent Tooling
- π§© Open-source agent tooling β MCP servers and CLI tools that let AI agents work with mail, messaging, notes, CI, and repos under least-privilege access controls; distributed via Homebrew and PyPI with Rust-accelerated local search.
- π Multi-agent research & synthetic-data pipelines β specialized agents research, analyze, and audit outputs before final generation, coordinated over a Rust/Tokio pub-sub event bus.
- π‘ Secure multimodal document agent β structured extraction from text and embedded images while treating document content as untrusted input, with prompt-injection-resistant validation.
- πΈ Cost-aware long-running agent harness β routes routine work to cheaper models, escalates hard work to stronger ones, and verifies outputs automatically to keep spend and reliability in balance.
Building AI products that stay fast, grounded, observable, and safe in production.





