- Design and build Auvik's event-driven alerting platform — alert generation, triggering, notification delivery, dismissal/clearing, and read-model updates across distributed Go services
- Own alert reporting and analytics: turning raw alert streams into severity analysis, resolution metrics, baselines, and trend reporting for managed service providers
- Led development of a modernized alert notification lifecycle system — a Kafka-backed, collection-based state machine that replaced ad hoc alert handling with an auditable event-driven lifecycle
- Ship backend services and APIs in Go across gRPC, GraphQL, REST, and Protocol Buffers, then wire them into GraphQL resolvers and React/TypeScript UI so alert data reads correctly end-to-end
- On-call debugging across Kafka consumers, Kubernetes/Helm workloads, Datadog metrics, and persisted collection data — usually the one tracing a bad alert back to the event that caused it
- Prototype internal AI tooling with Python, FastAPI, LangGraph, and OpenAI APIs, including a Slack-integrated agent over vector search
Problem: the legacy alert event pipeline fired alerts as fire-and-forget messages — no durable state, so knowing whether an alert was still active, already dismissed, or auto-cleared meant re-deriving it from scattered service logs. That made resolution metrics and trend reporting unreliable.
Design: the modernized system models each alert as a stateful entity backed by a Kafka-driven event log. Triggering, dismissal, clearing, and read-model updates are all separate events consumed by dedicated Go services, with collection-based state management persisting the current status instead of inferring it after the fact.
Backward compatibility: the legacy pipeline was still in production for other consumers, so dismissal and auto-clear logic had to stay correct for both architectures during the migration — no dual-write bugs, no alerts silently stuck open.
Surface: the improved read models flow through gRPC/GraphQL services into resolvers consumed by the React/TypeScript alerting UI, and into the alert reporting layer for severity, resolution-time, and baseline/trend analysis MSPs use to prioritize what to fix first.
Operating it: when something misbehaves in production, the trail runs through Kafka consumer lag, Kubernetes/Helm rollout state, Datadog dashboards, and the underlying Postgres/ClickHouse/Redis-backed collections — that's usually where the actual root cause shows up.