jj.dev
AVAILABLE · OPEN TO OPPORTUNITIES

Justin Johnson

Senior DevOps / Platform Engineer ★ AWS Certified Solutions Architect

Senior DevOps / Platform Engineer with 5+ years building cloud-native infrastructure on AWS and GCP, backed by a prior career leading large-scale enterprise operations across multiple lines of business. I bring both the technical depth to architect production platforms and the operational leadership context (cross-functional coordination, SLA governance, 900+ agents) to understand exactly what's at stake when they go down.

5+years cloud/DevOps
900+agents managed
30%IaC cut
25%deploy time saved
20%compute cost cut
jj@prod-cluster:~
$

Outcomes that matter

Numbers from real systems, real teams, real stakes.

~30%
Provisioning overhead cutReusable Terraform & Ansible modules across AWS and GCP
~25%
Deployment time reducedStandardized Jenkins CI/CD with Docker & Helm
~30%
Deployment failures downAutomated pipelines shipping to production multiple times/week
~20%
Compute costs savedOptimized container scheduling and workload allocation on EKS
5+ hrs
Saved per sprintInternal developer tooling that standardized environment setup
900+
Agents directedLarge-scale enterprise operations across multiple lines of business
+30%
Cross-sell close rateTargeted training programs for a 176-agent customer service team

Featured Work

Real deployed systems — production patterns, not tutorials.

GitHub
GH Actions
OIDC Auth
API Gateway
Lambda
DynamoDB
CloudFront
// AWS Solutions Architecture + DevOps

AWS Hybrid Architecture Portfolio

★ — stars — commits

A set of deployed, production-pattern AWS systems demonstrating Solutions Architect–level design decisions alongside DevOps operational rigor. Each module includes architecture decision records (ADRs) documenting the trade-offs made.

4
arch modules
0
long-lived credentials
ADR
per module
100%
IaC provisioned
01 · Serverless API 02 · Event-Driven Architecture 03 · Static Site + CDN 04 · CI/CD with OIDC ADR Decision Records
  • Serverless REST API — API Gateway, Lambda, DynamoDB
  • Event-driven system with SNS, SQS, dead-letter queues
  • Secure static delivery: S3 + CloudFront, private origins
  • GitHub Actions OIDC — zero long-lived AWS credentials
  • Architecture decisions documented per module
  • Deployed and tested against real AWS services
Docker
Ollama
RAG / LangChain
FastAPI Agent
Prometheus
React UI
// MLOps / LLMOps Platform Engineering

LLMOps Reference Platform

6 stages LLMOps platform

End-to-end LLMOps platform built to production deployment standards: model serving, RAG pipelines with vector search, CI/CD-gated evaluation, Prometheus observability, and multi-agent orchestration — deployed on containerized infrastructure reflecting enterprise AI patterns.

6
MLOps stages
0
external API calls
Live
WebSocket UI
Eval
CI pipeline
Stage 1 · Model Serving Stage 2 · RAG Pipeline Stage 2.5 · DevOps AI Agent Stage 3 · CI/CD for AI Stage 4 · Observability Stage 5–6 · Multi-Agent UI Security Hardening
  • Full LLMOps lifecycle: model serving, RAG, CI/CD evaluation, and observability
  • Zero external API dependency — local inference via Ollama with MLflow experiment tracking
  • DevOps AI agent with Docker socket access for container orchestration via natural language
  • CI/CD evaluation gate: automated pytest eval suite runs on every model or prompt change
  • Production observability: Prometheus, Grafana, and Evidently AI for model drift monitoring
  • Multi-agent orchestration with streaming React/WebSocket UI

Live Infrastructure Metrics

Real-time simulation of a production Kubernetes cluster — the kind of environment I build and operate.

CPU Utilization LIVE
64%
3/5 nodes active · EKS prod-cluster
Memory Usage LIVE
71%
11.4 / 16 GiB allocated
Deploy Success Rate 30d
98.2%
147 deployments · 3 rollbacks
Container Health LIVE
12/12
All pods running · 0 restarts (1h)
Pipeline Queue JENKINS
api-service / main ✓ PASS
worker / deploy-prod ● RUNNING
infra / terraform-plan ✓ PASS
ml-pipeline / retrain ● QUEUED
Uptime LIVE
99.97%
Last incident: 47 days ago

Cloud Architecture Explorer

Interactive diagram of the AWS Hybrid Architecture Portfolio. Click any service to learn more.

Career Timeline

Senior DevOps Engineer · CallTek
Oct 2024 – Present
  • Designed Jenkins CI/CD pipelines with Docker and Helm across dev, staging, and production. −25% deploy time
  • Built reusable Terraform & Ansible modules for AWS and GCP multi-region environments. −30% provisioning effort
  • Automated backup and disaster recovery workflows, improving RTO and platform resilience against unplanned outages.
  • Optimized container resource allocation and workload scheduling on managed EKS clusters. −20% compute cost
Freelance Software & DevOps Engineer · Independent
Oct 2023 – Oct 2024
  • Delivered cloud infrastructure, full-stack apps, and DevOps tooling for multiple clients, consistently on schedule.
  • Shipped React and Node.js applications end-to-end: UI, API integration, and AWS deployment with automated CI/CD.
  • Built and maintained personal AWS projects at production-grade reliability using containerization and automated pipelines.
  • Configured AWS-native services including S3, CloudFront, Lambda, and API Gateway for client production workloads.
DevOps / Software Engineer · Bell Media
Mar 2022 – Oct 2023
  • Implemented automated CI/CD pipelines enabling multiple production releases per week. −30% deploy failures
  • Built React components and Node.js REST APIs for internal editorial and content-ops tooling.
  • Created internal developer tooling that standardized environment setup and deployment. 5+ hrs saved/sprint
  • Embedded vulnerability scanning and secrets detection into delivery pipelines with the security team.
Senior Operations Manager · Bell Canada
Feb 2021 – Mar 2022
  • Directed operations spanning 900+ agents across multiple lines of business, maintaining SLA targets for enterprise service delivery. 900+ agents
  • Identified bottlenecks across tier-1 support workflows using data analytics and reporting. +15% SLA attainment
  • Led cross-functional incident response coordination, reducing mean time to resolution across business units.
  • Aligned operational reporting with ISO 27001 governance requirements across multiple business units.
Operations Manager · Bell Canada
Apr 2020 – Feb 2021
  • Managed 176 agents and implemented targeted training programs. +30% cross-sell close rate
  • Deployed KPI reporting dashboards that measurably reduced undetected SLA breaches across service lines.
  • Optimized staffing allocation models and shift workflows, reducing overtime costs.
  • Worked with IT to implement proactive monitoring alerts for earlier issue detection before customer impact.

Technical Toolkit

Hover to explore. Every tool here has production mileage behind it.

AWS EKS IAM VPC CloudWatch GCP GKE Docker Helm Linux RHEL Terraform Ansible Jenkins GitHub Actions CI/CD Pipelines Kubernetes Container Scheduling Prometheus Grafana ELK Stack Alerting Image Scanning Secret Detection AWS Secrets Manager Parameter Store OIDC Python Bash / Shell Groovy JavaScript
Cloud & Containers
IaC & Orchestration
Observability
Security
Languages

Credentials

🏅
AWS Certified Solutions Architect — Associate
Amazon Web Services
🎓
Software Development
Bell University
⌨️
Full Stack Engineering
Codecademy
📊
Project Management · Business Ops · Data Analytics
Microsoft Learn