Site Reliability Engineer · Kochava
Aug 2024 — Present · Remote
- Operate hybrid multi-cloud infrastructure for cross-cloud products and services — AWS and GCP alongside on-prem, including Kubernetes clusters that span all three environments.
- Own the cross-cloud network fabric: site-to-site VPN tunnel architecture, routing, and connectivity between AWS, GCP, and on-prem.
- Deploy and run self-hosted platform services end to end (Airflow, databases) — provisioning, upgrades, and reliability — where self-hosting beats managed offerings on cost or control.
- Drive FinOps across both clouds: cost optimization folded into architecture decisions up front, not retrofitted after the bill arrives.
- Own reliability, monitoring, and observability across cloud and on-prem systems; land defaults that are secure, compliant, and cost-effective at the same time.
- Partner with ML teams on MLOps infrastructure: Dask clusters for distributed compute and Airflow-orchestrated model-training pipelines.
- Build platform automation — from infrastructure workflows to onboarding scripts that turn manual setup into repeatable, auditable runs.
- Embed AI-led development practices and workflows into infrastructure engineering, working across cross-region teams.