Projects

RL environments and agent evals, backed by the distributed-systems work that makes them realistic.

DevOps Gym

RL Environments

A suite of Dockerized, reproducible DevOps tasks for training and evaluating AI agents. Each task ships a realistic broken state (failing CI pipeline, broken K8s deployment, crashing service), a task spec, a hidden test-based verifier, and a difficulty tier — the building blocks of RL-environment work.

Key Features

  • Dockerized, reproducible task sandboxes
  • Hidden test-based verifiers per task
  • Difficulty tiers and task specs
  • Reward-hacking-resistant grading
DockerPythonBashKubernetesGitHub Actions

K8s Debug Evals

Evals

A reproducible eval harness for measuring frontier models on Kubernetes debugging. Runs agents against broken clusters, scores pass@1 with hidden verifiers, and classifies failures into a taxonomy — built to demonstrate rigorous, non-vibe-coded eval methodology.

Key Features

  • pass@1 scoring against hidden verifiers
  • Failure taxonomy across models
  • Reward-hacking detection examples
  • Fully reproducible runs via Docker + kind
PythonKuberneteskindClaude APIOpenAI API

K8s Job Scheduler

Distributed Systems

A Go-based HTTP API server for prioritizing and submitting Kubernetes jobs. Features a max-heap priority queue, concurrency control, and graceful shutdown handling — the orchestration depth that makes my DevOps environments realistic.

Key Features

  • Priority queue with max-heap implementation
  • Kubernetes client-go integration
  • Concurrency control mechanisms
  • Graceful shutdown handling
GoKubernetesclient-goREST APIMax-Heap

HybridRAG

LLM Systems

A cost-optimized enterprise search engine with a three-layer architecture: Ingestion Pipeline for async document processing, Hybrid Retrieval combining semantic and keyword search with reranking, and a Cost Router for intelligent model selection — LLM-systems work end to end.

Key Features

  • Async document ingestion pipeline
  • Semantic + keyword search with reranking
  • Cost-aware model routing
  • Production-ready architecture
PythonFastAPINext.js 14Qdrant CloudCohereGroq Llama-3

Vector Database

Fundamentals

A complete vector database implementation from scratch in Python. Supports multiple distance metrics (Euclidean, Cosine, Dot Product), CRUD operations, persistence, and semantic search — fundamentals, built by hand.

Key Features

  • Multiple distance metrics
  • Top-k similarity search
  • Metadata storage and retrieval
  • Semantic search integration
PythonNumPySentence Transformers

Why environments, not demos?

The known failure mode of the RL-environment market is vibe-coded tasks with graders that agents can trick. Everything here is built the other way: reproducible sandboxes, hidden verifiers, explicit difficulty calibration, and documented reward-hacking resistance. Building systems from scratch — vector databases, schedulers, RAG pipelines — is how I learned what rigor looks like.