Site Reliability Engineer · Talon.One
Shan E Raza
I keep large systems observable, reliable, and—on the good days—boring.
Berlin-based. A decade of metrics, logs and traces—and disaster recovery as a reflex, not a fire drill.
What I work on
- ObservabilitySurfacing problems across the fleet before they become incidents.
- Platform & GitOpsSelf-service platforms—engineering teams provision their own infra, treated as customers.
- Database resilienceKeeping Postgres fast and right-sized under the loads that actually bite.
- Critical user journeysDecomposing distributed architectures into the flows that matter, then the health checks and SLOs that measure what people actually feel.
Operating creed
- If it can break, it will—so design for the failure, not the demo.
- You build it, you run it. Ownership beats a handoff, every time.
Stack
PrometheusThanosGrafanaLokiOpenTelemetryKubernetesGKEAWSGCPAzureArgoCDTerraformHelmCrossplanePostgresincident.ioGoPythonBash