Years Exp
Prod Nodes
SLO Attained
I build calm systems for chaotic problems. Agents that explain themselves, pipelines that heal themselves, and infrastructure that lets people sleep.
If a task happens twice, it becomes a script. If the script runs twice, it becomes a pipeline.
An agent nobody can audit is a liability. Every tool call and model decision leaves a span in the trace.
A working demo in the customer's environment beats a hundred slides. Prototype first, polish under load.
Years in Production
Nodes Managed
MTTR Reduced
Incidents Documented
Architected high availability EKS clusters with Karpenter for rapid node provisioning. Reduced cloud costs by 25% via Spot instances.
Bedrock PoC: Strands Agents on AgentCore Runtime that investigate AWS infra anomalies autonomously. Every tool call and model decision traced via OTEL/ADOT into CloudWatch.
Agentic ITSM workflows in ServiceNow: AI driven ticket triage, incident enrichment from telemetry, and automated resolution suggestions with human in the loop approval.
Modular Terraform library for VPC, RDS, and ECS provisioning. Enforced policy-as-code using OPA/Checkov in CI pipelines.
Enterprise retrieval pipeline on Bedrock Knowledge Bases with OpenSearch vectors. Chunking strategies tuned per document type, citations returned with every answer.
Multi cluster EKS delivery with ArgoCD and Kustomize. Every environment reconciled from git, drift detected and reverted automatically within minutes.
A latent race condition in DynamoDB's DNS automation wiped the regional endpoint's records, triggering a 15 hour cascade across the world's busiest cloud region.
READ_LOG →Scattered Spider vished an outsourced helpdesk into resetting MFA, stole Active Directory, and ransomed a FTSE 100 retailer's ESXi estate.
READ_LOG →Chinese state hackers compromised nine US telecoms and reached the CALEA lawful intercept systems built for court-authorized wiretaps.
READ_LOG →
Amazon Web Services
Microsoft
Anthropic
Professional Certification
Embed with your team. Map the workflow, the data, and the pain before writing a line of code.
A working agent in your environment within days. Real data, real integrations, real feedback.
Guardrails, evals, and observability. Every decision traced, every failure mode rehearsed.
Docs, training, and a clean path to production. Your team owns it when I leave.
I sit with the customer, not behind a ticket queue. FDEs take an ambiguous business problem, embed with the team that owns it, and build a working solution in their real environment. Less handoff, more shipping.
That is the day job. Recent examples include ServiceNow ticket triage agents and an autonomous observability agent on Amazon Bedrock. Reach out with the problem, not the solution, and we will scope it together.
AWS first: Bedrock, AgentCore, Strands Agents, EKS, Terraform. Observability through OpenTelemetry into CloudWatch and Grafana. For ITSM work, ServiceNow. Pragmatic about everything else.
The blog hosts 70+ deep dives on outages and breaches, from CrowdStrike to Salt Typhoon. Each covers the timeline, the technical root cause, and the lessons. Visit bl0g.surge.sh.