The brief
Accelerate development of aligned, safe, and powerful AI by building the highest-quality human-in-the-loop data infrastructure at scale — encompassing annotator training and certification, multi-project annotation execution across multiple AI lab clients, containerized reproducible dataset environments, and the internal tooling needed to manage annotation workforce quality and throughput at industrial scale.
What we built
A full-stack AI data infrastructure platform serving as the operational backbone for human-in-the-loop model alignment work. Phase 1 (Oct–Dec 2025): Annotator onboarding on Outlier platform, RLHF/SFT training data generation for Claude Sonnet, TTS model annotation (Guitar Pinstripe/Riff), and UI annotation (Meadow UI). Phase 2 (Mar–May 2026): Industrial-scale Dockerized dataset registry (5,000+ Docker images across multiple programming languages pushed to AWS ECR), large-scale agent trajectory generation for Claude and Gemini models (Jaeger/Jaeger 2.0 projects), ARC-AGI-style reasoning game suite (Arc Agents), VINDEX response quality evaluation, KAIJU repository validation, Leviathan PRD generation for web-building AI agents, and an ETP (Ethara Task Platform) internal workflow tool with rubric-based QC flows. Throughout: production EKS infrastructure with GPU/CPU Karpenter autoscaling, Istio service mesh, multi-model AWS Bedrock hosting (Claude, Qwen, Kimi K2.5, MiniMax, GLM), and full GitOps CI/CD via ArgoCD and GitHub Actions.



