Brian Lopez
+1 (818) 298-0568 |
bloinlagr@gmail.com |
linkedin.com/in/thebrianlopez |
brianlopez.us
Chicago, IL (Remote-ready, travel available)
Summary
Platform engineer operating at staff scope, with 15+ years across infrastructure automation, site reliability, and internal platform engineering - including SRE management at AXS before returning to hands-on IC work by choice to build at IPO scale. Joined Grindr (NYSE: GRND) pre-IPO and helped carry the infrastructure through its NYSE listing and 126% revenue growth ($195M to $440M). Currently responsible for multi-account AWS infrastructure (8 AWS accounts, 5+ EKS clusters) serving 15M MAU and 300M+ daily messages across 190 countries: IaC, GitOps via ArgoCD, CI/CD pipeline automation, Datadog observability, and cross-system integrations that permanently remove manual operator steps. Builds internal developer tooling in Go (workctl, mdq, perfgate) adopted by engineering teams. Also operates AI/ML inference infrastructure and an event-driven telemetry pipeline that measures AI-assisted engineering workflows. Bilingual (English / Spanish).
Skills
Infrastructure & Cloud:
- IaC: Terraform (multi-account, module authoring, state management, Atlantis PR-gated plan/apply), Helm, ArgoCD (GitOps, app-of-apps, upgrade pipelines)
- Cloud: AWS (EKS, EC2, RDS, MSK, S3, IAM, VPC, KMS, Secrets Manager, CloudFront, Route53, ACM, SES, DynamoDB, EMR, Firehose, CodeDeploy, ElastiCache)
- Kubernetes: Multi-cluster EKS operations, Karpenter, cluster-autoscaler, external-dns, cert-manager, CSI drivers, IRSA, PodDisruptionBudgets
- Networking: VPC peering (multi-account), cross-account IAM, Cloudflare, AWS Load Balancers, Traefik, Consul
Observability & Reliability:
- Datadog (SLOs, dashboards, alerting, K8s orchestration view), PagerDuty, Prometheus, Grafana
- On-call design, blameless post-mortems, incident management
Languages & Tooling:
- Go (CLIs, internal platform tooling, test suites, telemetry instrumentation)
- Python, Bash/Fish shell scripting
- GitHub Actions, CI/CD pipeline design
Data & ML Platform:
- Kafka / Confluent Schema Registry, Apache Airflow (Astronomer), Snowflake, Databricks, EMR
- LiteLLM (centralized LLM proxy), ML inference infrastructure (training, decider, enricher services)
Agentic Engineering / LLMOps:
- Agent orchestration: multi-agent task dispatch with an explicit state machine, milestone coordination, and idempotent hook-based triggers
- Telemetry for AI-assisted engineering: schema-versioned event bus with session correlation, cost attribution, and runaway-loop detection, feeding a human-reviewed rule-update loop
Experience
Grindr
- Own Terraform infrastructure across 8 AWS accounts supporting a 15M MAU platform in 190 countries: account-level responsibility for VPC architecture, IAM auth flows, cross-account peering, and secrets management
- Maintain OneLogin/AWS SSO federation across all 8 accounts via Terraform: SAML role topology (admin/readonly tier separation, ExternalId-gated role enumeration), OAuth client credential lifecycle in Secrets Manager, and S3 + SQS + Lambda audit log ingestion pipeline feeding the SIEM
- Operate ArgoCD GitOps across 5+ EKS clusters, managing upgrade branches, app-of-apps patterns, and CRD lifecycle for all environments
- Designed and built Grindr's alpha environment from scratch: VPC, EKS cluster, external-dns, cert-manager, and 10+ production services deployed via GitOps
- Own infrastructure for Musubi, Grindr's AI-powered content moderation platform: training service, inference API, decider service, and data enricher running in dev and prod; designed database migration safety gates (reader/writer endpoint topology assertions, schema validation, alembic smoke tests in deploy pipeline) after a production DDL incident
- Automated Databricks infrastructure CI/CD via GitHub Actions and eliminated static credential risk across all Databricks environments: provisioned scoped IAM for a CI service account to execute auditable, repeatable deployments autonomously; migrated authentication from long-lived PAT tokens to Service Principal OAuth (M2M), removing the platform team from the data engineering deploy critical path permanently
- Deployed LiteLLM proxy for centralized LLM cost tracking and governance across engineering teams; own ongoing security posture (CVE patching, version pinning, internal-network exposure controls)
- Built multi-VPC private connectivity between Chalk ML platform and production MSK: secure cross-account ML inference path
- Operate cross-repo GitHub Actions automation for SOX-bounded analytics workloads: keep dbt projects in sync between primary and SOX-segregated repos via PR-gated GitHub Apps
- Designed and maintains cross-system automation integrating Jira, GitHub, Confluence, and an internal JSONL event bus: code-reviewed, tested, and continuously measured; eliminates manual workflow fragmentation across the engineering organization (
workctl platform)
- Built internal developer tooling in Go:
workctl CLI (Jira + GitHub + Confluence integration), mdq (markdown query engine), perfgate (statistical performance gates)
- Led systematic RDS cost-reduction campaign: traffic analysis showed sustained low utilization on 5 production read replicas and 2 oversized instances; removed and right-sized them incrementally, delivering a six-figure monthly reduction in database spend with zero read-latency regressions
- Drove proactive infrastructure security hardening: identified and remediated internal credentials embedded in git history across infrastructure repos; migrated Hazelcast license key from static config to AWS Secrets Manager via External Secrets Operator; enforced private networking for AI platform endpoints before production traffic traversed public paths
Joined pre-IPO; supported Grindr through its NYSE listing (November 2022, ticker: GRND). Platform scaled from 12M to 15M MAU and 111B+ annual messages (300M+/day) during this period.
- Completed migration of PB-scale data lake from EC2-based Cloudera cluster to Snowflake: phased EOL, data validation, zero data loss
- Supported Chat3 replatform: server-side migration of message storage and logic to sustain 300M+ daily messages across 15M MAU with end-to-end encryption
- Deployed Confluent Schema Registry for polyglot Kafka support across the data platform; deployed Kafka topics and MSK access patterns for 10+ microservices
- Designed and maintained Datadog SLOs, dashboards, and alerting for Kubernetes workloads tied to user-visible degradation, not infra noise; participated in monthly on-call rotation with documented post-incident reviews
- Provisioned and operated microservices via Helm, ArgoCD, GitHub Actions, and AWS CodeDeploy with auto-scaling and self-healing across dev/prod
- Conducted IAM security hardening: remediated identity-based MSK policy, provisioned Secrets Manager across dev/prod environments
- Wrote and maintained automation tooling in Go, Python, and Bash; managed IaC with Terraform across EC2, RDS, EKS, S3, KMS, CloudFront, and Route53
AXS
- Led SRE for AEG's live event ticketing platform across 30+ major venue clients and 9 operating regions (US, UK, JP, SE, CA, AU, NZ, FR, DE), including NFL/NBA/NHL/concert properties and Japan's 36-team B.League deployment
- Owned reliability for a dual-stack platform - AXS Global (PHP/Symfony on AWS ECS, Node.js Unified API on Elastic Beanstalk) plus legacy Veritix Mainline (.NET/C# VAScheduler) - maintaining one SLA through the COVID-19 mass event cancellation period
- Built production SRE automation: CloudFront cache invalidation for high-demand onsales, PostgreSQL snapshot-restore, Elasticsearch reindex workflows, and Lambda event sourcing with multi-region IAM across US and JP
- Operated centralized observability and edge protection: Sumo Logic via AWS FireLens/Fluent Bit, CloudWatch alarms, and Cloudflare Enterprise WAF/bot management for waiting room and Unified API traffic
- Ran blameless post-mortems; reduced major incident recurrence; established on-call rotations, escalation paths, and runbook culture across L1–L3 engineers
- Established SLO framework across the ticketing platform - defined reliability targets for systems supporting live event transactions across US, UK, and Japan; gave teams and leadership a shared data language for reliability investment
- Drove gradual, validated rollout patterns for production change propagation across high-risk platform surfaces (checkout, ticket delivery, waiting room onsale flows)
- Contributed to AWS infrastructure across the platform estate: EC2, Elastic Beanstalk, ECS, S3, CloudFront, Kinesis, Lambda, DynamoDB, Secrets Manager
- Supported production environment reliability; drove uptime improvements through incident analysis, tooling, and platform operations
- Mentored colleagues on DevOps practices; contributed to on-call discipline and runbook culture
- Managed production environment across DataCenter technologies and AWS (EC2, Elastic Beanstalk, RDS, CloudFront, ElastiCache)
- Responsible for production support, on-call, and DevOps tooling
Nestlé
- Developed and deployed global client management solution for 250,000 systems worldwide: Microsoft System Center, PowerShell, .NET Framework
- Built ASP.NET web service and PowerShell modules automating image build process (SCCM, MDT, Active Directory)
- Improved service delivery via operational KPIs from infrastructure telemetry, help-desk tickets, and monitoring
- Automated day-to-day operations using PowerShell modules and .NET applications; ran L1–L3 training workshops
- Level 3 escalation engineer for North America; collaborated with operating companies globally
- Managed 30,000 Windows systems: security patches, software standards, configuration parity
- Maintained global knowledge base; ran daily/weekly operational review meetings
Pomeroy
- Level 2 escalation engineer supporting 14,000 business systems across factory, office, and HQ locations
- Subject matter expert on business integration projects from acquisitions (Gerber, Kraft Pizza, Dryers)
Self-Employed
- Configured and deployed business solutions for SOHO/SMB clients using Microsoft and Linux/Unix technologies
- Managed Active Directory, file servers, web servers, VPNs within agreed SLAs
Education
California State University, Northridge - Coursework, Electrical and Computer Engineering
2001 – 2005
College of the Canyons - CCNA Coursework, Computer Systems Networking
2007 – 2009
Projects
workctl - Go CLI integrating Jira, GitHub, Confluence, and a custom JSONL event bus into a unified engineering workflow tool. Full test coverage and structured telemetry: internal tooling built to remove friction from how engineers work.
Runabout Devtools (Go) - Suite of six engineering utilities: mdq (markdown query engine over markdown frontmatter and table cells, used for cross-document reporting), perfgate (statistical performance gates for CI:distribution comparison vs. point estimates), shellprof (shell-function latency profiler that drove fish startup optimization), hookval, wasend, protonexport. Each with more than 100 tests and telemetry instrumentation that feeds the Automation Metrics event bus.
Automation Metrics: Telemetry for AI-Assisted Engineering - Event-driven observability system for agent-assisted development workflows: a schema-versioned event bus with session correlation spanning editor hooks, shell, Go CLIs, and LLM inference. Runs a closed feedback loop: hooks emit events, a measurement pipeline compares behavior against baselines, proposed rule updates pass human review, and the behavior change is re-measured. Caught and corrected a model-routing cost regression and detects runaway agent loops before they burn spend. Single-operator system built with production observability discipline.
Farewell John: Multi-Cloud Infrastructure for an Early-Stage Startup - Independent infrastructure ownership for hersolutionsllc (4-person startup). Multi-cloud Terraform IaC across AWS (Cognito + Lambda auth triggers, CloudFront/S3, Amplify, Route53/ACM, SES, Secrets Manager) and GCP (Cloud Run, Firebase, IAM service accounts, multi-project topology). Operates 7 Apache Airflow DAGs for cross-system automation: auth validation monitoring, ETL pipelines, geo-pricing crawlers, data backfills. Email deliverability hygiene (Zoho Mail with MX/SPF/DMARC). Administer the custom-domain Google Workspace tenant. Production application with real users.
Multi-Agent Epic Delivery System - Coordination substrate for multi-agent software delivery. Markdown + YAML frontmatter dispatch files with a strict three-state machine (pending to claimed to complete) enforced by SessionStart and UserPromptSubmit hooks. Supports both implementation dispatches (Type 1, with epic milestone tracking) and discovery dispatches (Type 2, with write-back convention). Sentinel-file pattern prevents duplicate event emissions across prompts. Coordinates concurrent releases across multiple projects with a test-first discipline layer and a written post-mortem for every agent failure mode.
Linkari: Adaptive Personal AI Assistant - Mobile-first agentic UX evolving from a score-and-forget link curator into an adaptive personal assistant. Three strategic pillars in flight: (1) implicit-feedback loop that converts digest interactions into rubric-weight calibration signals, (2) semantic clustering of shares for emergent topic tracking, (3) action routing that drafts Jira tickets, digest entries, or pins based on score + profile. Self-hosted single-user architecture using the Claude Code CLI via OAuth2 (no API keys).
Idea Workflow Engine (POC) - GitHub Actions + Claude pipeline that moves startup ideas through Spark, Research, POC, and PRD stages with state checkpointing, GitHub Issue-driven clarifying questions, and automatic artifact linking between idea, research doc, and downstream PRD. Demonstrates agent-native workflow orchestration where the LLM is the host process and CI is the actuator.