The stack
The tech stack we build, run and trust in production.
Model-agnostic. Tool-rich. Observed end-to-end. We compose the right layer per workflow, and bring guardrails, evals and human oversight from day one.
Stack layers
Models
Frontier and open models, picked per task to balance quality, latency, cost and safety.
- OpenAI GPT-5 / o-series
- Anthropic Claude Sonnet & Opus
- Google Gemini 2.x
- Meta Llama, Mistral, Qwen (self-hosted)
- Whisper, Deepgram, ElevenLabs (speech)
Agent frameworks
Graph-based orchestration with explicit state, tools, retries and human checkpoints.
- LangGraph & LangChain
- Mastra, CrewAI, AutoGen
- OpenAI Agents SDK
- Vercel AI SDK
- Custom orchestrators when frameworks get in the way
Tools & integrations
Typed, permissioned tools that let agents act, not just answer.
- MCP servers (custom + open source)
- Function calling / structured outputs
- EHR & FHIR connectors
- Twilio, Slack, Gmail, Calendar, Stripe
- Browser automation (Playwright, Browserbase)
Retrieval & memory
Grounded answers with hybrid retrieval, citations and per-document access control.
- pgvector, Pinecone, Weaviate, Qdrant
- Hybrid BM25 + dense retrieval
- Reranking (Cohere, Voyage)
- Long-term memory stores (Mem0, Letta)
- Document parsing: Unstructured, LlamaParse
Realtime voice
Sub-second voice agents with barge-in, function calling and human warm-transfer.
- OpenAI Realtime, Gemini Live
- LiveKit Agents, Pipecat, Vapi
- Twilio Voice, Telnyx
- Deepgram & Whisper STT
- ElevenLabs & Cartesia TTS
Evals & observability
Continuous evaluation and tracing ensure agents are tested like systems, not vibes.
- LangSmith, Langfuse, Braintrust
- Helicone, Arize, OpenTelemetry
- Golden datasets + LLM-as-judge
- Regression and drift alerting
- Per-trace cost & latency budgets
Security & compliance
BAA-covered providers, PHI minimization and tenant isolation by default.
- BAAs with OpenAI, Anthropic, AWS, GCP
- PHI redaction & data minimization
- RBAC, audit logs, tenant isolation
- SSO / SCIM, KMS-managed secrets
- HIPAA + SOC2-aligned controls
Workflow & data
Pipelines, queues and warehouses that feed agents and absorb their outputs.
- Temporal, Inngest, Trigger.dev
- Postgres, Supabase, Redis
- BigQuery, Snowflake, dbt
- Kafka & event streams
- Airbyte / Fivetran connectors
Infra & deployment
Edge-fast, GPU-ready and cost-tuned to run wherever your data is allowed to live.
- AWS (Bedrock, ECS, Lambda)
- GCP Vertex AI, Azure OpenAI
- Cloudflare Workers / Vercel edge
- Modal, Replicate, RunPod for GPUs
- Docker, Terraform, GitHub Actions
The engineering and ML stack that hosts the agents.
Agents don't live in a vacuum. They sit inside real apps, services and data platforms, and not every workload needs an LLM. Here's the rest of what we build and run on.
Product engineering

React

Node.js

Next.js

Angular

Express

Vue

JavaScript

TypeScript

Sequelize

React Native

Flutter

.NET Core

Figma

WordPress

AWS

PostgreSQL

MySQL

MongoDB
Classical ML & data science
TensorFlow
PyTorch
SciPy
Pandas
NumPy
OpenCV
Scikit-learn
Theano
Keras
NLTK
SpaCy
ChromaDB
