Senior DevOps Engineer
Alpaca · Remote - Japan - APAC
Job description
About the role
As a Senior DevOps Engineer you will design, build and operate the infrastructure that lets Alpaca scale globally and run trading‑critical systems with confidence. You will have autonomy to design and implement solutions against clearly defined goals and a real voice in shaping those goals with the team.
Key responsibilities
- Design and evolve cloud architecture on GCP (networking, IAM, high‑availability topology) expressed entirely as Terraform code with GitOps.
- Build and own CI/CD pipelines for IaC, including policy‑as‑code guardrails, drift detection and progressive rollouts.
- Advance Platform‑as‑Product by creating self‑serve capabilities and golden paths for engineers.
- Strengthen the observability stack (Prometheus, Thanos, Grafana, Loki, Tempo, Alertmanager) for metrics, logs, traces and alerting.
- Operate GKE clusters and associated services (Helm‑packaged workloads, RabbitMQ, IBM MQ, PostgreSQL).
- Participate in a Follow‑The‑Sun on‑call model: triage alerts, lead incidents, conduct blameless post‑mortems and drive corrective actions.
- Embed SRE practices (SLIs/SLOs, error budgets, capacity planning) into core infrastructure development.
Required profile
- 5+ years in DevOps, Platform/Infrastructure, or SRE roles operating large‑scale, high‑availability systems.
- Deep hands‑on experience designing GCP cloud architecture (landing zones, networking, IAM, HA).
- Strong Infrastructure‑as‑Code expertise with Terraform, large multi‑environment codebases, and a GitOps mindset.
- Proven ability to build CI/CD pipelines for IaC with automated plan/apply, code review and policy enforcement.
- Production experience with Kubernetes (preferably GKE) and Helm deployments.
- Solid networking fundamentals (VPC, routing, load balancing, DNS, TLS) and ability to debug cross‑service connectivity.
- Hands‑on experience with modern observability tools (Prometheus, Thanos, Grafana, Loki, Tempo, Alertmanager).
- Operator‑level familiarity with PostgreSQL and message brokers such as RabbitMQ or IBM MQ.
- Good understanding of SRE practices (SLIs, SLOs, error budgets) and a Platform‑as‑Product mindset.
- Strong incident management skills, including participation in APAC on‑call rotations and effective async communication.
Required skills
- Google Cloud Platform (GCP)
- Terraform
- GitOps
- CI/CD pipelines for IaC
- Kubernetes / GKE
- Helm
- Prometheus, Thanos, Grafana, Loki, Tempo, Alertmanager
- PostgreSQL
- RabbitMQ / IBM MQ
- Advanced cloud networking (VPC, routing, load balancing, DNS, TLS)
What we offer
- Competitive salary and stock options
- Health benefits
- New hire home‑office setup (one‑time USD $500)
- Monthly stipend of USD $150 via a Brex card
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Japan.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 9 hours ago
Expires 1 month from now
4 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Alpaca
Remote - Japan - APAC