Senior SRE / DevOps Engineer (Kubernetes / A2P Messaging)
This environment will suit you if you care not only about whether infrastructure is running, but about how the application behaves on top of it.
We usually respond within a week
This environment will suit you if you enjoy building infrastructure from the ground up, solving hard reliability problems and staying close to the systems you operate in production.
You’ll build and operate the infrastructure behind a new high-volume messaging platform, taking ownership from the Kubernetes cluster and networking layer through observability, deployment and incident response.
The platform processes around 1 million messages every day, supports approximately 175 customers and 120 suppliers, and maintains more than 300 long-lived messaging connections.
This is not a typical stateless web environment. The platform runs long-lived TCP sessions that need stable network endpoints and predictable behaviour during deployments, failures and traffic spikes.
You’ll help build the platform across two infrastructure sites and stay with it through production readiness, migration, hypercare and steady-state operation.
Rate: 150–190 PLN + VAT (B2B) per hour.
Location: Fully remote (Poland) 🌎 or Łódź (Poland)

Your everyday challenges:
Use AI-assisted engineering where it improves delivery - including infrastructure tooling, automation, documentation and runbooks.
Build and operate Kubernetes infrastructure across two sites using infrastructure-as-code and automated delivery.
Design deployment mechanisms that allow long-lived connections to drain gracefully instead of being dropped during releases.
Take ownership of platform networking, including stable ingress and egress IPs, Layer 4 load balancing, TLS and connectivity coordinated with an external infrastructure provider.
Build and operate the observability stack: metrics, logs, dashboards and alerting that provide useful signal rather than noise.
Design monitoring around real production behaviour, including visibility into individual connections and failure modes.
Own certificate lifecycle, secrets management and platform access control.
Operate PostgreSQL in a highly available environment, including replication, failover, backups and verified restores.
Design and exercise backup and disaster-recovery procedures across two sites.
Work closely with engineers and QA on performance and production-scale load testing, investigating what fails first and why.
Support the platform during migration and production cutovers.
Take part in production incident response and from the steady-state phase, a 24×7 on-call rotation.
What are we looking for?
Experience with telecom, carrier, messaging or other environments built around persistent network connections.
Knowledge of SMPP, SIP, SS7 or similar telecom protocols.
Strong production experience operating Kubernetes on self-managed, on-premise or similarly infrastructure-heavy environments - not only managed cloud services.
Experience with stateful, long-lived TCP workloads on Kubernetes, including connection draining, stable ingress/egress, Layer 4 load balancing and deployment behaviour.
Strong Linux and networking fundamentals. You are comfortable diagnosing problems involving routing, NAT, firewalls, MTU, TLS or packet-level behaviour.
Practical troubleshooting experience with tools such as tcpdump and production network diagnostics.
Experience designing monitoring and alerting, not only maintaining dashboards somebody else created. You understand what deserves to wake a human up, and what does not.
Strong infrastructure-as-code and CI/CD experience for containerised systems.
Practical PostgreSQL operations experience, including replication, failover, backup and importantly - verified restore.
Ability to work directly with engineers from external infrastructure and network providers.
Professional English.
Real production on-call and incident-response experience.
What will strengthen your candidacy?
Experience configuring or troubleshooting IPsec connectivity with external parties.
Experience designing or operating multi-site active-active or active-passive environments.
Hands-on disaster-recovery exercises rather than DR plans that existed only on paper.
Performance engineering experience, including Linux kernel or network tuning for high connection counts.
Security hardening experience, including CIS-style benchmarks, vulnerability management or software supply-chain practices.
Experience being the first SRE or Platform Engineer on a system and defining how it should be operated.
Why join us?
Build it and run it: you won’t inherit an infrastructure somebody else designed or throw your work over the wall after launch. You’ll help build the platform and remain close to it in production.
A genuinely difficult reliability problem: long-lived protocol traffic, stateful connections, two-site infrastructure and production traffic at telecom scale.
Influence architecture from day one: deployment, observability and operability are design constraints here, not tasks postponed until after development.
Greenfield infrastructure: you’ll have real influence over how the Kubernetes platform, delivery pipelines, monitoring and operational practices are created.
Small senior team: short decision paths and direct collaboration with engineers and the solution architect.
Production ownership: you’ll follow the system through build, migration, hypercare and steady-state operation instead of disappearing after implementation.
Modern engineering environment: automation and AI-assisted tooling are used where they genuinely improve engineering and operational work.
Sounds like a fit? Let’s talk! 🚀
- Locations
- Łódź (Monopolis)
- Remote status
- Fully Remote
- Employment type
- Full-time
Łódź (Monopolis)
Why work with us
-
You may hear it frequently from leaders at DNA that they would like to attract people smarter than themselves. Leaders at DNA feel responsible for creating an environment in which self-motivated people build teams that are able to create great things.
-
The source of productivity comes from a team having clearly defined autonomy and responsibility, not from micro-managing people.
-
We believe that diversity and psychological safety at work matter — we get proofs of that everyday.
-
Usually, the problems we are trying to solve are complex (not complicated). Frequently, it means that we follow the Agile process and build our organisation according to the DevOps model.
Join a team with great autonomy and responsibility. The team discusses business priorities with the Product Owner, defines architecture, technology stack, way of working, takes pride in engineering craftsmanship and is end-to-end responsible for the solution.
About DNA Technology
As part of the Digital New Agency (DNA) group, DNA Technology is an experienced technology partner supporting clients in solving complex problems using software. DNA Technology is a long-term travel companion to startups and industry leaders. DNA offices are located in Stockholm and Łódź.