Signal, not noise

Paulina Stopa

Unix Engineer · AI-Driven Infrastructure & SRE Automation

I run reliability across Sky's critical Unix estate by day and operate my own autonomous AI-operations platform around the clock — LLM agents, RAG-backed runbooks and gated fleet automation, all governed like production. Nine years of enterprise SRE discipline, now pointed at what operations becomes next.

01 — Profile

Reliability by day, autonomy by night

Paulina Stopa

Nine years running Sky's critical estate — a fleet of roughly 10,000 physical and virtual RHEL hosts managed with Ansible, AWX on a K3S Kubernetes cluster our team designed and deployed, Red Hat Satellite and corporate LDAP. I fulfil senior-level responsibilities across reliability, automation and identity — and the team's internal tooling portfolio, from AWX to the LDAP management tools and the Redirect Platform, carries my work throughout.

My exposure runs the full stack: from assembling the hardware, through operating systems, virtualisation, storage, monitoring and security, to deploying and developing the tools on top — finishing at SEO and marketing.

In my own time I build and operate a complete autonomous-operations platform on my own infrastructure: LLM agents working 24/7 behind human review gates, RAG-backed runbook automation, self-hosted model serving and Prometheus/Grafana observability on the AI platform. It is run like production — versioned, monitored, governed — because that is the only way operations automation earns trust.

Before Sky, I founded and ran an eight-person company. I bring SRE discipline, working systems that prove the approach, and a track record of making AI adoption real.

~10,000
hosts in Sky's critical Unix estate
24/7
autonomous LLM agents, behind human review gates
63 / ~120
containers on one GPU VM / across the whole estate
20 + 12
Ansible playbooks & roles driving self-healing fleet automation
02 — Track record

Experience

Unix Engineer — Sky UK Ltd

Leeds, United Kingdom · February 2017 – Present · previously Shift Leader (2018–2019)

Engineer in Sky's largest Unix team, fulfilling senior-level responsibilities — reliability, automation and engineering standards across a ~10,000-host critical estate of physical and virtual RHEL systems (plus legacy Solaris) underpinning national-scale services.

  • Fleet automation at scale: production automation in Python, Bash and Ansible across the estate — provisioning, patching and lifecycle via Red Hat Satellite, orchestration through AWX running on a K3S Kubernetes cluster our team designed and deployed.
  • Identity engineering: operate the corporate 389-ds LDAP estate powering authentication, groups, netgroups, sudo policy and automount fleet-wide; led the data conversion and migration of the central LDAP service authenticating tens of thousands of users — plus third-party retail LDAPs and legacy ODSEE systems maintained for service continuity.
  • Internal tooling portfolio: helped design, build and maintain the team's tool estate — LDAPtool and uID (identity management web tools), the Redirect Platform (hundreds of Sky domains including sky.com and skysports.com), GoAccess log analytics, Redash compliance reporting (Lynis, MariaDB cluster), PCS cluster status for Pacemaker/Corosync HA, a Caddy reverse proxy with ACME certificates, an MkDocs knowledge base and the Unix Tools Dashboard.
  • Integrations & workflow automation: built integrations between CyberArk, GitHub and MS Teams, and automated email notification workflows for operational events.
  • Standards & enablement: own documentation and engineering standards for the wider team (Backline function); train and onboard engineers to production standard; former Shift Leader with Incident Coordinator training — hands-on incident management from detection through resolution and review.
  • AI adoption: drive practical AI use across Sky's engineering community — a regular contributor on GenAI, containerisation and security, with work recognised by the Head of Enterprise Product Strategy (GenAI).

Independent AI Engineering — Autonomous Operations Platform

Self-directed engineering programme · October 2024 – Present

Designed, built and operate a production-grade autonomous-operations estate on my own infrastructure — the full stack from hardware assembly through virtualisation to AI services. Everything below runs today: versioned as code, monitored, governed.

  • Sky Unix AI Platform (concept & vision demo): built solo, on this platform and in my own time, an interactive mock-up exploring AI-assisted operations for a Unix estate — LLM incident triage over a runbook knowledge base, remediation behind human approval gates and workflow visualisations — plus a costed requirements analysis of the path from proof-of-concept to production. Built to test feasibility and inspire AI adoption at Sky.
  • AI platform: 63 Docker containers on a GPU-passthrough Proxmox VM (2× NVIDIA P100) — LiteLLM model gateway, Ollama local inference (Gemma, Llama, Nemotron, Qwen), a self-hosted Dify RAG stack with Qdrant, and speech, image, video and music generation services behind Traefik ingress; all versioned Docker Compose infrastructure-as-code in self-hosted Gitea with CI. Evaluated and self-hosted NVIDIA DGX Spark (2026).
  • Autonomous agents: two LLM agent harnesses running 24/7 on dedicated service accounts behind human review gates — executing scheduled operational tasks across 27 skill categories and authoring incident and postmortem reports published as internal sites.
  • ClawBoard (author; open source, MIT): agent-operations dashboard with real-time agent status over WebSocket, a full transcript audit trail, JWT authentication and a gated production/development deployment workflow — github.com/Wadera/clawboard.
  • Fleet & hosting estate: 20 Ansible playbooks and 12 roles executed via self-hosted Semaphore across a two-node Proxmox HA cluster (~120 containers estate-wide) running Debian, Ubuntu, RHEL, Rocky and Windows — snapshot-before-change with automated restore, per-host-class patching and Vault-encrypted secrets; plus a classic hosting layer operated for years: DirectAdmin, Proxmox Mail Gateway, Prometheus, Uptime Kuma, Gitea, Nextcloud and production WordPress, Ghost, Drupal and e-commerce (WooCommerce, PrestaShop) sites.

Founder & Owner — SKYDAY

Chojnów, Poland · April 2012 – February 2017

Built one of Poland's largest voice-server platforms — 10,000+ concurrent users and 100,000+ registered accounts — including the custom CRM and billing system behind it.

  • Led a company of eight as founder and managing director — budgeting, client acquisition, government grant funding, and end-to-end project delivery for clients including the Polish National Bank and Play (mobile operator).
  • Operated dedicated-server infrastructure hosting hundreds of domains and services — from hardware to SEO and marketing delivery.

Co-founder — Digital Technologies S.C.

Poland · November 2014 – November 2016

Among Poland's first certified commercial UAV operators — aerial photogrammetry and 3D measurement models for construction and solar-installation clients.

03 — Capabilities

Skills

SRE & Reliability

SLIs / SLOs & error budgets Incident management Observability Capacity planning Toil reduction Change management Self-healing automation

AI & LLM Operations

LLM integration (LiteLLM, Ollama) Agentic automation Retrieval-augmented generation Self-hosted inference Human-in-the-loop governance Prompt engineering

Automation & Infrastructure as Code

Python Bash Ansible (AWX, Semaphore) AWX on K3S Kubernetes Red Hat Satellite podman Docker & Compose CI/CD (Gitea CI) GitOps Proxmox virtualisation

Cloud & Edge

OVH Hetzner IONOS Cloudflare WAF / DDoS / CDN K3S Kubernetes VMs & storage

Observability & Telemetry

Prometheus Grafana Uptime Kuma Alerting & noise reduction Dashboard design Telemetry-informed automation

Security & Identity

389-ds LDAP CyberArk integration Least-privilege design Secrets management (Vault) Cyber security OSINT Audit trails for AI output
04 — Recognition

Recognition

Cyber Champions Black Badge

First person at Sky to achieve the highest status in the company-wide security initiative.

Sky Stars

Multiple nominations, including recognition from the Head of Enterprise Product Strategy (GenAI).

05 — Credentials

Certifications & Training

Incident Coordinator

Sky operational-readiness training — hands-on incident management from detection through resolution and review.

Ansible & Kubernetes Training

Formal training in fleet automation and container orchestration.

OSINT Certifications

Certified in open-source intelligence techniques.

UAV Operator Certificate

25 kg class · Civil Aviation Authority, 2014.

Education

BEng Computer Science: Systems and Networks

Witelon University of Applied Sciences in Legnica, Poland · 2010–2014

Languages

Polish — native · English — professional working proficiency

Italian — elementary

06 — Off duty

Beyond the terminal

Entrepreneur at heart — before Sky I founded and ran an eight-person company end to end, from budgeting and client acquisition to delivery. When the laptop closes I am usually on a motorcycle, somewhere with more horizon than signal. Relentlessly curious and open-minded, I build things — servers, platforms, businesses — and want to understand how everything, and everyone, works.

07 — Contact

Get in touch

Open to conversations about AI-driven infrastructure, SRE and platform engineering — the running estate and the platform demos are the best evidence I can offer, and I am happy to walk you through both.