praca hybrydowa

Lokalizacja:

Warsaw

Data dodania:

2026-08-31

Wymagane umiejętności:

Site Reliability Engineering (SRE)
Kubernetes
CI/CD pipelines
Prometheus
IaC Terraform
Linux
Puppet
Hashicorp Vault
Bash
Python
ELK
Ansible
Grafana

Mile widziane:

Forma zatrudnienia:

Opis projektu:

The Technology Infrastructure Services Engineer will be responsible for designing, implementing, operating, and continuously improving enterprise-scale Ceph storage platforms that support critical business and research workloads across hybrid and cloud-native environments. The primary focus of the role is to build, support, automate, and operationalize Ceph infrastructure, ensuring high availability, scalability, performance, and resilience. The engineer will lead operational improvements by developing automation, infrastructure-as-code, monitoring, and self-service capabilities while driving standardization and reducing operational overhead. In addition to Ceph, the engineer will have opportunities to work with complementary storage technologies such as Weka and Qumulo, along with Kubernetes, HashiCorp Vault, and modern infrastructure platforms. This role requires a strong DevOps and Site Reliability Engineering (SRE) mindset, with responsibility for improving platform reliability through automation, observability, proactive monitoring, incident response, capacity planning, and continuous improvement. The ideal candidate enjoys operating large-scale distributed systems and is passionate about eliminating manual work through engineering and automation.

Twój zakres obowiązków:

  • Design, implement, operate, and continuously improve enterprise-scale Ceph storage platforms supporting critical business and research workloads across hybrid and cloud-native environments.
  • Build, support, automate, and operationalize Ceph infrastructure, ensuring high availability, scalability, performance, and resilience.
  • Lead operational improvements by developing automation, infrastructure-as-code, monitoring, and self-service capabilities.
  • Drive standardization and reduce operational overhead.
  • Work with complementary storage technologies such as Weka and Qumulo, Kubernetes, HashiCorp Vault, and modern infrastructure platforms.
  • Improve platform reliability through automation, observability, proactive monitoring, incident response, capacity planning, and continuous improvement.
  • Operate large-scale distributed systems and eliminate manual work through engineering and automation.

Nasze wymagania:

  • Minimum 2 years’ experience.
  • Hands-on experience designing, deploying, administering, and supporting Ceph storage clusters in production environments.
  • Experience operationalizing Ceph, including lifecycle management, upgrades, expansion, health monitoring, capacity planning, troubleshooting, and performance tuning.
  • Experience automating Ceph operations using Ansible, Terraform, Python, Bash, or similar automation frameworks.
  • Experience building operational tooling, runbooks, and self-service capabilities to improve platform efficiency and reliability.
  • Strong understanding of distributed storage concepts, including replication, erasure coding, CRUSH maps, OSDs, MONs, MGRs, CephFS, RBD, and RGW.
  • Experience integrating Ceph with Kubernetes or cloud-native platforms is highly desirable.
  • Experience with Kubernetes and containerized infrastructure.
  • Knowledge of HashiCorp Vault or similar secrets management solutions.
  • Experience with Linux platform administration.
  • Understanding of networking concepts supporting distributed storage platforms.
  • Strong understanding of Site Reliability Engineering (SRE) principles.
  • Experience of using Puppet tools for configuration management.
  • Experience implementing Infrastructure-as-Code using Terraform and/or Ansible.
  • Experience building CI/CD pipelines for infrastructure automation.
  • Experience with monitoring, alerting, logging, and observability platforms (Prometheus, Grafana, ELK, etc.).
  • Experience with incident management, root cause analysis, and operational excellence.
  • Strong scripting skills (Python, Bash, Go, or similar).
  • Experience of other HPC storage platforms will be a bonus.
  • Experience supporting hybrid cloud infrastructure (AWS preferred).
  • Experience operating storage platforms supporting AI/ML, HPC, or large-scale Kubernetes environments.

To oferujemy:

  • Ubezpieczenie na życie
  • Spotkania integracyjne

Benefity:

Prywatna opieka medyczna

Ubezpieczenie na życie

Spotkania integracyjne

Dofinansowanie zajęć sportowych

Interesuje Cię ta oferta?

Sprawdź podobne oferty:

Analityk
praca hybrydowa

Lokalizacja:

Warszawa

Data dodania:

2026-09-09

PM/PMO
praca hybrydowa

Lokalizacja:

Warszawa

Data dodania:

2026-09-09

praca hybrydowa

Lokalizacja:

Warsaw, Gdansk

Data dodania:

2026-09-09

Lokalizacja:

Warsaw, Gdańsk, Gdynia

Data dodania:

2026-09-09

praca hybrydowa

Lokalizacja:

Warsaw

Data dodania:

2026-09-08