MANAGER OF CLOUD SYSTEMS · THERAPYNOTES

Taylor
Finklea

Leading teams. Building systems. Shipping outcomes.

Leads SRE and DevOps at TherapyNotes: architecting the Kubernetes platform, enabling teams to own its implementation and operations, commanding major incidents, and building the tools and practices behind safe, practical AI adoption.

tip: press ⌘K to jump anywhere

system status ● all operational
Engineering Leadership
teams own outcomes
Platform Engineering & Reliability
reliable production operations
Applied AI Engineering
safe, practical adoption
// 01

Teams, platforms, and applied AI

Engineering leadership, reliable platforms, and applied AI work together to create durable outcomes.

01

Engineering Leadership

Building autonomous teams, developing technical leaders, commanding major incidents, and aligning engineering work with business outcomes.

02

Platform Engineering & Reliability

Architecting cloud and Kubernetes platforms, improving delivery systems, and enabling teams to own reliable production operations.

03

Applied AI Engineering

Building AI tools, dashboards, standards, and workflows that make adoption practical, safe, and effective across engineering.

// 02

How I lead

I build teams that own outcomes. That means creating clear responsibility, developing technical leaders, removing processes that do not add value, and staying close enough to the work to provide meaningful technical direction.

01

Build teams around the work

Formed and staffed an SRE team around strong incident responders, then reshaped responsibilities across engineering groups so ownership better matched each team’s strengths. Developed leads and managers so execution could remain team-owned rather than manager-driven.

02

Lead calmly through incidents

Serve as Incident Commander during major production events. I bring in the right responders, protect engineers from unnecessary interruption, and keep leadership and customer-facing stakeholders informed while technical teams restore service.

03

Replace process theater with ownership

Moved teams away from hour-based estimation, time tracking, and manager-controlled sprint mechanics. Put delivery and process improvement into the hands of team leads and engineers, using small experiments and observation instead of sweeping methodology changes.

04

Stay technical where it creates leverage

Architected the Kubernetes platform and helped establish its initial direction before transitioning implementation and operations to the team. Remain hands-on in applied AI by building tools, dashboards, standards, and practices that help engineering teams adopt AI safely and effectively.

// 03

Selected work

Showing 12 of 14 projects
008 Applied AI Engineering Published

Harness Deck

One review surface across every coding harness

Built a harness-neutral local dashboard where AI coding agents publish durable reports, mockups, approvals, and decisions through a versioned contract.

Claude CodeOpenAI CodexOpenCodePi coding agentModel Context ProtocolHomebrew
009 Applied AI Engineering Published

Browser Agents

vision-driven web automation

AI-powered browser agents that navigate and automate complex web workflows: identifying elements by appearance and context instead of brittle selectors, with retries, waits and audit trails.

SkyvernStagehandPlaywrightFirecrawl
010 Applied AI Engineering Published

AI Development Environment

AI as an everyday tool

Built the environment and guardrails that made AI a safe, everyday tool for engineering: approved models, local LLMs, evals and observability, and a culture excited to share AI-enabled work.

GitHub CodespacesOllamaLM StudioMCPLLM evaluation
011 Applied AI Engineering Published

AI Automation Agents

reason · plan · execute

Production AI agents that reason, plan and execute multi-step work: customer comms, lead research/enrichment, email and form automation, and API orchestration. Python-first for reliability.

LangGraphPydantic AIMCPReActLangSmith
012 Platform Engineering & Reliability Published

Kubernetes

architect and technical lead

Architect and technical lead for the Kubernetes platform. Evaluated Docker Swarm and Nomad, chose Kubernetes and Helm, and helped establish both managed (AKS) and immutable self-managed (Talos, RKE2 on Micro OS) directions before transitioning implementation and operations to the team.

KubernetesTalosRKE2AKSHelm
013 Applied AI Engineering Published

Making AI Safe and Practical Across Engineering

From isolated experiments to everyday engineering practice

Developed standards, tools, dashboards, and a community of practice that helped teams adopt AI safely while preserving engineering judgment and accountability.

Azure AI FoundryAnthropic Claude PlatformOpenAI PlatformModel quota managementHITRUST AI readinessAI communities of practice
015 Engineering Leadership Published

Process That Earns Its Place

Autonomy inside real guardrails

Evaluated each team practice on the value it delivers, removed time tracking and manager-run sprint mechanics, and grew team autonomy while keeping governance and compliance boundaries intact.

LeanAgileKataChange management
016 Engineering Leadership Published

Leading Major Incidents Without Becoming the Bottleneck

Clear command, protected responders

Serves as Incident Commander during major events: bringing in the right people, protecting technical responders from noise, and keeping stakeholders informed while teams restore service.

Incident commandSREObservability
017 Platform Engineering & Reliability Published

Zero Trust

smaller attack surface

One of the most complex problems I've owned: Cloudflare Zero Trust and a bastion for access, a hub-and-spoke firewall managed with GitOps, and a move toward ACL-based access with Tailscale.

CloudflareAzure FirewallTailscaleGitOps
018 Platform Engineering & Reliability Published

Automation & CI/CD

self-service pipelines

Built CI/CD on Azure DevOps then migrated to GitHub Actions: reusable workflows, OIDC service principals, KeyVault secrets, and Automation Accounts for patching and scripting.

Azure DevOpsGitHub ActionsOIDCKeyVault
019 Platform Engineering & Reliability Published

Infrastructure as Code

immutable, rebuilt at will

Replaced hand-run PowerShell with Terraform and Packer: custom reusable modules, ephemeral systems, and persistent user data via FSLogix roaming profiles on Windows.

TerraformPackerHCLFSLogix
020 Platform Engineering & Reliability Published

Containerize & Monitor

no more 3am pages

Containerized manually-deployed microservices, built a custom CI process with Ansible and Bash, and stood up Prometheus / Grafana / Alertmanager monitoring that was still in use when I left.

DockerPrometheusGrafanaAnsibleCI
021 Platform Engineering & Reliability Published

Migrating to Azure

Over 200 VMs · zero downtime

Planned and executed the migration of two datacenters (over 200 VMs from ESXi to Azure) with zero downtime, plus a cloud-native SIEM to replace legacy on-prem.

AzureSentinelMigrationO365
// 04

Experience

A decade of growing scope, from frontline support to leading teams, architecting platforms, and building applied AI systems.

Manager of Cloud Systems

2023 - Present
TherapyNotes

Leads SRE and DevOps: architecting the Kubernetes platform, enabling teams to own its implementation and operations, commanding major incidents, and building the tools and practices behind safe, practical AI adoption.

// 05

Capabilities

Engineering Leadership

  • Organizational change
  • Cross-functional collaboration
  • Vendor management
  • Cost management
  • Lean
  • Agile
  • SRE
  • Kata
  • Change management
  • Strategic planning
  • Conflict resolution
  • Active listening
  • Project management
  • Cloud engineering management
  • IT management

AI Platforms and Enterprise Adoption

  • Anthropic Claude Platform
  • Anthropic API
  • ChatGPT Enterprise administration
  • OpenAI Enterprise management
  • Enterprise AI governance
  • Production model enablement
  • Model quota management
  • HITRUST AI readiness
  • AI adoption strategy
  • AI communities of practice

Agent Systems and Developer Tooling

  • Model Context Protocol
  • MCP
  • Agent Skills
  • Claude Code skills
  • Codex skills
  • OpenCode skills
  • Pi extensions
  • Ralph loops
  • /loops
  • beads
  • Harness Deck
  • Conductor
  • Human-in-the-loop workflows
  • Multi-model routing
  • Context engineering
  • Durable agent handoffs
  • Structured AI documentation
  • Agent verification workflows
  • A2A
  • AI agent development
  • Homebrew

Applied AI Engineering

  • Stagehand
  • Skyvern
  • Ollama
  • LM Studio
  • Local LLMs
  • LLM evaluation
  • AI observability
  • LangSmith
  • Langfuse
  • LangWatch
  • Prompt engineering
  • Tool calling
  • Structured outputs
  • OpenAI
  • Claude
  • Gemini
  • ReAct
  • Groq
  • faster-whisper
  • Piper
  • ElevenLabs

Platform Engineering and Reliability

  • Terraform
  • OpenTofu
  • Packer
  • Ansible
  • Docker
  • GitHub Actions
  • Azure DevOps
  • CI/CD
  • Spacelift
  • Prometheus
  • Grafana
  • Datadog
  • Tailscale
  • Infrastructure as code
  • DevSecOps
  • Incident response
  • Observability
  • GitHub Codespaces
  • Git

Cloud, Security, and Compliance

  • Active Directory
  • Linux
  • Windows Server
  • SOC 2 Type II
  • ISO 27001
  • HITRUST
  • HIPAA
  • NIST
  • OWASP
  • Tenable.io
  • Nessus
  • Microsoft Sentinel
  • Microsoft Defender
  • Privileged Identity Management
  • PIM
  • Change Advisory Board
  • CAB
  • Jamf
  • Jamf Protect
  • Mimecast
  • SecurityScorecard
  • IT audit
  • Vulnerability scanning
  • Penetration testing
  • Dependabot
  • Kali
  • Intune
  • Ubiquiti
  • Juniper

Software Development

  • Bash
  • PostgreSQL
  • SQLite
  • InfluxDB
  • MySQL
  • SvelteKit
  • Tauri
  • SwiftUI
  • FastAPI
  • Vue
  • Nuxt
  • Supabase
  • Firebase
  • Cloudflare Workers
  • Cloudflare Pages
  • Vercel
  • Fly.io
  • Railway
  • Netlify
  • Loro CRDT
  • FTS
  • iCloud
  • App locking
  • Portable export
  • Figma

// WHAT'S NEXT

Let’s build what’s next.

I work at the intersection of principal engineering and senior leadership: building teams that own outcomes, architecting reliable platforms, and making AI practical across engineering. If that’s the leader, architect, or applied AI engineer your organization needs, let’s talk.

BEYOND WORK

  • Home lab: Self-hosted infrastructure and Home Assistant.
  • Wildlife photography: Birds, insects, and the natural world.
  • Birding and naturalism: Exploring the land around my home near Kansas City.
  • Daily hiking: Time outside, every day.
Taylor Finklea
esc