Signal, not noise

Paulina Stopa

Unix Engineer · AI-Driven Infrastructure & SRE Automation

I run reliability across Sky's critical Unix estate by day and operate my own autonomous AI-operations platform around the clock — LLM agents, RAG-backed runbooks and gated fleet automation, all governed like production. Nine years of enterprise SRE discipline, now pointed at what operations becomes next.

01 — Profile

Reliability by day, autonomy by night

Paulina Stopa

Nine years engineering reliability on Sky's critical Unix estate — roughly 10,000 hosts underpinning national-scale services. The day job is classic SRE: production automation in Python, Bash and Ansible that removes toil, standardises change and keeps a large, complex estate dependably boring.

Outside it, I build and operate a complete autonomous AI-operations platform on my own infrastructure: LLM agents working 24/7 behind human review gates, RAG-backed runbook automation, self-healing fleet automation with snapshot-and-restore, and Prometheus/Grafana observability on the AI platform, with automated monitoring maintenance windows across the fleet. It is run like production — versioned, monitored, governed — because that is the only way operations automation earns trust.

At Sky, I designed and demoed the Sky Unix AI Platform — an AIOps proposal for LLM incident triage with human-gated remediation. At home, I operate its working counterpart every day.

~10,000
hosts in Sky's critical Unix estate
24/7
autonomous LLM agents, behind human review gates
63 / ~120
containers on one GPU VM / across the whole estate
20 + 12
Ansible playbooks & roles driving self-healing fleet automation
02 — Track record

Experience

Unix Engineer — Sky UK Ltd

Leeds, United Kingdom · February 2017 – Present · previously Shift Leader (2018–2019)

Nine years in Sky's largest Unix team, fulfilling senior-level responsibilities — delivering reliability, automation and engineering standards across a ~10,000-host estate behind national-scale services.

  • Sky Unix AI Platform: designed and demoed an end-to-end AIOps proposal for the Unix estate — ServiceNow incident ingestion, LLM triage and root-cause analysis over a RAG runbook knowledge base, and AWX/Ansible remediation behind human approval gates — delivered as four interactive demo modules plus a costed requirements report.
  • Develop and maintain production automation in Python, Bash and Ansible across the estate — eliminating repetitive toil, standardising change and reducing operational risk.
  • Led the migration of the central LDAP service authenticating tens of thousands of users — a critical availability dependency for the entire estate.
  • Mentored engineers through the Redirect platform migration — configuration standards and design review across hundreds of Sky domains, including sky.com and skysports.com.
  • Former Shift Leader with Incident Coordinator training — hands-on incident management on a national-scale platform, from detection through resolution and review.
  • Own documentation and engineering standards for the wider team; drive practical AI adoption across Sky's engineering community, recognised by the Head of Enterprise Product Strategy (GenAI).

Independent AI Engineering — Autonomous Operations Platform

Self-directed engineering programme · 2024 – Present

Designed, built and operate a production-grade autonomous-operations estate on my own infrastructure — the working counterpart to the platform proposed at Sky. Everything below runs today: versioned as code, monitored, and governed.

  • AI platform: 63 Docker containers on a GPU-passthrough VM (~120 containers estate-wide) — LiteLLM model gateway, Ollama self-hosted inference, a Dify RAG stack with Qdrant vector search, and speech services behind Traefik ingress; all versioned Docker Compose infrastructure-as-code with CI.
  • Autonomous agents: two LLM agent harnesses running 24/7 on dedicated service accounts behind human review gates — executing scheduled operational tasks and authoring incident and postmortem reports.
  • ClawBoard (author; open source, MIT): agent-operations dashboard with real-time agent status over WebSocket, a full transcript audit trail, JWT authentication and a gated production/development deployment workflow — github.com/Wadera/clawboard.
  • Fleet automation: 20 Ansible playbooks and 12 roles run via self-hosted Semaphore across a two-node Proxmox HA cluster — snapshot-before-change with automated restore, per-host-class patching, automatic monitoring maintenance windows and Vault-encrypted secrets.
  • Observability: Prometheus and Grafana on the AI platform, with automated monitoring maintenance windows across the fleet to keep signal fidelity high during change work.
  • Evaluated and self-hosted NVIDIA DGX Spark for local AI workloads (2026).

Founder & Owner — SKYDAY

Chojnów, Poland · April 2012 – February 2017

Built one of Poland's largest voice-server platforms — 10,000+ concurrent users and 100,000+ registered accounts — with a custom CRM and billing system, on self-operated dedicated-server infrastructure serving clients including the Polish National Bank and Play.

Co-founder — Digital Technologies S.C.

Poland · November 2014 – November 2016

Among Poland's first certified commercial UAV operators — aerial photogrammetry and 3D measurement models for construction and solar-installation clients.

03 — Capabilities

Skills

SRE & Reliability

SLIs / SLOs & error budgets Incident management Observability Capacity planning Toil reduction Change management Self-healing automation

AI & LLM Operations

LLM integration (Claude, LiteLLM, Ollama) Agentic automation Retrieval-augmented generation Self-hosted inference Human-in-the-loop governance Prompt engineering

Automation & Infrastructure as Code

Python Bash Ansible (AWX, Semaphore) Docker & Compose CI/CD (Gitea CI) GitOps Proxmox virtualisation

Cloud & Edge

OVH Hetzner IONOS Cloudflare WAF / DDoS / CDN Kubernetes concepts VMs & storage

Observability & Telemetry

Prometheus Grafana Alerting & noise reduction Dashboard design Telemetry-informed automation

Security & Identity

LDAP / IAM Least-privilege design Secrets management (Vault) Cyber security OSINT Audit trails for AI output
04 — Recognition

Recognition

Cyber Champions Black Badge

First person at Sky to achieve the highest status in the company-wide security initiative.

Sky Stars

Multiple nominations, including recognition from the Head of Enterprise Product Strategy (GenAI).

05 — Credentials

Certifications & Training

Incident Coordinator

Sky operational-readiness training — hands-on incident management from detection through resolution and review.

Ansible & Kubernetes Training

Formal training in fleet automation and container orchestration.

OSINT Certifications

Certified in open-source intelligence techniques.

UAV Operator Certificate

25 kg class · Civil Aviation Authority, 2014.

Education

BEng Computer Science: Systems and Networks

Witelon University of Applied Sciences in Legnica, Poland · 2010–2014

Languages

Polish — native · English — professional working proficiency

Italian — elementary

06 — Contact

Get in touch

Open to conversations about AI-driven infrastructure, SRE and platform engineering — the running estate and the platform demos are the best evidence I can offer, and I am happy to walk you through both.