Skip to content
View SteamyCutie's full-sized avatar
πŸ’‘
Good ideas !
πŸ’‘
Good ideas !
  • HOME-BASED

Organizations

@NVIDIAGameWorks @onpointtech @ThetaStash @eLearningDAO

Block or report SteamyCutie

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SteamyCutie/README.md

AI/ML-Focused Full Stack Software Engineer

typing banner

profile views focus open to collaboration


πŸ™‹β€β™‚οΈ About Me

Β Β Full stack software engineer with 10+ years of experience shipping production systems, now centered on AI-powered products: retrieval, LLM applications, recommendation systems, and agentic workflows.

Β Β I own products end-to-end β€” UI, APIs, data pipelines, model serving, cloud infrastructure, observability, and AI evaluation. Work has spanned systems serving millions of interactions, with 30%+ latency reductions and 2Γ—+ inference throughput gains through model routing, batching, caching, and autoscaling.

Β Β I care about the unglamorous parts that make AI usable in production: groundedness and evaluation, cost per request, tracing and observability, and guardrails against prompt injection and tool abuse.


🎯 Focus Areas

Area What I build
πŸ€– Agentic AI Multi-step agent loops, tool calling, structured outputs, MCP servers, human-in-the-loop approvals, long-running orchestration
πŸ”Ž RAG & Retrieval Hybrid search, embeddings, Graph-RAG, reranking, permissions-aware retrieval, long-context pipelines
⚑ LLM Serving & Cost Unified model gateways, dynamic routing, provider failover, vLLM/Ray Serve, batching, caching, warm model pools
πŸ“ˆ Search & Recommendations Learning-to-rank, personalization models, behavioral signals, real-time inference, A/B experimentation
πŸ›  Data & ML Platform Streaming/feature pipelines, embedding + indexing freshness, MLOps, deployment and rollout automation
πŸ§ͺ Evaluation & Observability Automated eval suites, hallucination/groundedness/tool-accuracy metrics, OpenTelemetry tracing, token & cost analytics
πŸ” AI Security Prompt-injection defense, untrusted-content isolation, least-privilege tool access, RBAC/IAM, secrets management
🌐 Full Stack Product Streaming UIs, SSR/ISR apps, REST/GraphQL/gRPC APIs, event-driven microservices, payments and webhooks

🧠 Core Skills

Languages

Python TypeScript JavaScript Java Go Rust C++ SQL Bash

AI / ML & GenAI Engineering

PyTorch scikit-learn Hugging Face LangChain LangGraph CrewAI OpenAI Agents SDK vLLM MCP pgvector FAISS MLflow

RAG Graph-RAG Embeddings & Vector Search Agentic Workflows Multi-Agent Systems Agent Loops & Harnesses Tool Calling Structured Outputs Context Engineering Prompt Engineering Multimodal AI Synthetic Data Generation Model Routing LLM Evaluation Fine-Tuning (LoRA/QLoRA) MLOps ML Pipelines Pydantic LangSmith

Backend & APIs

FastAPI Spring Boot Django Flask Node.js GraphQL gRPC Kafka RabbitMQ

REST WebSockets Server-Sent Events Microservices Serverless Event-Driven Architecture Distributed Systems Concurrency Caching System Design DDD SOLID Design Patterns

Frontend & Full Stack

React Next.js Tailwind CSS Redux Angular Supabase Stripe

App Router SSR/ISR Zustand Streaming UI Responsive Design Accessibility Webhooks

Data, Search & Streaming

PostgreSQL MySQL Redis MongoDB Elasticsearch ClickHouse Snowflake Spark Airflow

DynamoDB BigQuery ETL/ELT Streaming Pipelines Query Optimization Schema Design Feature Pipelines

Cloud & DevOps

AWS GCP Azure Kubernetes Docker Terraform GitHub Actions Jenkins Linux

EC2 ECS Lambda S3 RDS Bedrock SageMaker Helm Argo CD Ansible Infrastructure as Code CI/CD

Testing, Reliability & Observability

pytest Playwright Vitest Jest Prometheus Grafana OpenTelemetry Datadog

TDD Unit / Integration / E2E API Testing Coverage Gates Load & Performance Testing Automated Evaluation Pipelines Distributed Tracing SLOs/SLIs Incident Response

AI Security & Application Security

OWASP Top 10 RBAC IAM OAuth 2.0 SSO/SAML Secrets Management Least-Privilege Tool Access Prompt-Injection Defense Tool-Abuse Guardrails Untrusted-Content Isolation Secure Agent Tooling


🚧 Selected Work

  • 🧩 Open-source agent tooling β€” MCP servers and CLI tools that let AI agents work with mail, messaging, notes, CI, and repos under least-privilege access controls; distributed via Homebrew and PyPI with Rust-accelerated local search.
  • πŸ”€ Multi-agent research & synthetic-data pipelines β€” specialized agents research, analyze, and audit outputs before final generation, coordinated over a Rust/Tokio pub-sub event bus.
  • πŸ›‘ Secure multimodal document agent β€” structured extraction from text and embedded images while treating document content as untrusted input, with prompt-injection-resistant validation.
  • πŸ’Έ Cost-aware long-running agent harness β€” routes routine work to cheaper models, escalates hard work to stronger ones, and verifies outputs automatically to keep spend and reliability in balance.

πŸ“Š GitHub Stats

top languages by repo top languages by commit commit stats

contribution streak

contribution chart


Building AI products that stay fast, grounded, observable, and safe in production.

Pinned Loading

  1. cronaswap-interfacev2 cronaswap-interfacev2 Public

    TypeScript 12

  2. cronaswap-protocolv3 cronaswap-protocolv3 Public

    TypeScript 2

  3. pulsechainart-backend pulsechainart-backend Public

    JavaScript 1

  4. pulsechainart-frontend pulsechainart-frontend Public

    TypeScript 1

  5. spring-petclinic spring-petclinic Public

    Forked from spring-projects/spring-petclinic

    A sample Spring-based application

    CSS 1 1

  6. workingdogs workingdogs Public

    TypeScript