Justin Bailey / Software & Platform Engineering

I build platforms
people can depend on.

Software and platform engineer, technical lead, and U.S. military veteran. I build Kubernetes operators, self-service infrastructure, and the practices that support production reliability.

Professional work

Selected engineering work.

Full experience

BrainGu

Platform software and delivery

Built a Kubernetes operator for tenant onboarding, RBAC, network policy, secrets, and upgrade safety gates.

Platform delivery with developers, product teams, security engineers, and customers.

Explore platform lifecycle work

Striveworks

Production reliability

Led high-severity incidents through cross-functional diagnosis, service restoration, and post-incident remediation.

Production reliability for a multi-region, GPU-backed MLOps platform.

Explore incident leadership

GovCIO

Infrastructure automation

Built configuration and bare-metal deployment tooling with Ansible and Nautobot.

Enterprise and edge environments supporting mission-critical systems.

Role & responsibilities

Architecture / Implementation / Operation

Independent practice

Projects that show how I think, build, and maintain systems.

Infrastructure practice

Home operations

Ongoing personal work

A hands-on environment for operating Kubernetes, exploring infrastructure patterns, and learning from real maintenance work.

  • Talos
  • Flux
  • Cilium
  • Proxmox
Explore

What I bring

Skills in practice.

Each area connects to work I have contributed to.

01

Platform software

Turn infrastructure requirements into APIs, controllers, and reusable capabilities with clear ownership and predictable reconciliation.

  • Go
  • Kubernetes APIs
  • Crossplane
  • CRDs & operators
Kubernetes lifecycle work
02

Developer experience

Build self-service workflows that let teams provision infrastructure and deliver software with less manual coordination.

  • GitOps
  • Terraform
  • GitLab CI/CD
  • Service catalogs
Self-service infrastructure experience
03

Production reliability

Connect telemetry to decisions: service objectives, incident response, release readiness, and distributed systems troubleshooting.

  • OpenTelemetry
  • Grafana
  • SLOs
  • Incident response
Incident leadership work
04

Secure delivery

Bring security into the build and operating model through container maintenance, policy, and repeatable delivery workflows.

  • Container security
  • SBOMs
  • Policy as code
  • Automation
Container maintenance experience

Field notes

Engineering notes.

All writing