Platform & SRE Engineer
About The Company
MWDN is a global IT outstaffing company with 23+ years of experience that connects exceptional tech talent with leading companies across Israel, the USA, Great Britain, and Western Europe. We offer opportunities to work on international products in a stable and professional environment.
Why does MWDN rock?
Here’s what you can expect when you join MWDN:
- Security: We carefully vet our clients to minimize risks and ensure reliability and timely payments-no fraud or unpleasant surprises.
- Career support: If a project isn’t the right fit, we support you and actively help find new opportunities that match your skills and career goals.
- Legal assistance: We provide guidance on legal matters, including opening and managing your independent contractor or sole proprietorship status, taxes, and related processes.
- Professional development: We offer English courses and professional growth opportunities, as well as team-building events.
Why choose us? MWDN is ranked among the top 5 IT employers in our region according to DOU. We take pride in our transparency and strong commitment to our team. Curious to learn more? See what our employees say about working with us on DOU.
What is your new project?
Domain: AI Infrastructure / Real-Time Data Processing
Client Location: Israel
Company size: 10 - 51
A fast-growing product company building next-generation infrastructure for production AI systems. The team is focused on high-performance, low-latency distributed technologies that operate directly within real-time data flows and AI workloads. The product addresses complex engineering challenges related to runtime reliability, distributed processing, networking, observability, and large-scale system performance.
What makes this project exciting?
The client is looking for a Platform & SRE Engineer to join the engineering team and take ownership of the infrastructure and operational foundation behind our distributed runtime environments.
This is a hands-on engineering role combining Linux systems engineering, AWS infrastructure, machine-image management, deployment automation, observability, security hardening, and SRE practices.
A key part of the role is owning the lifecycle of Linux-based runtime hosts — from AMI creation and machine provisioning through operating-system configuration, system services, runtime deployment, host tuning, upgrades, security hardening, monitoring, and operational troubleshooting.
You will maintain and evolve our internal host management and deployment tooling, ensuring runtime environments can be deployed, configured, upgraded, diagnosed, and operated reliably across cloud and, in the future, customer-hosted environments.
In parallel, you will help establish and own our SRE and observability practices, ensuring we can measure platform health and availability, identify failures quickly, and operate against clearly defined reliability objectives.
What makes you a great fit
- 8+ years of experience in DevOps, SRE, Platform Engineering, Linux Systems Engineering, or Infrastructure Engineering.
- Strong Linux systems administration and troubleshooting skills, including systemd, processes, networking, filesystems, permissions, kernel configuration, and system-level debugging.
- Strong hands-on AWS experience, particularly with EC2, networking, IAM, storage, and production Linux workloads.
- Strong experience with AWS AMIs, including building, configuring, hardening, validating, versioning, and maintaining machine images for production environments.
- Strong experience with Linux server hardening and production security practices, including least privilege, service isolation, user and permission management, secure system configuration, and attack-surface reduction.
- Hands-on experience with Grafana, OpenTelemetry, Prometheus, Mimir,
- ClickHouse, Fluent Bit, or comparable observability technologies.
- Experience building production monitoring, dashboards, metrics, logging, and alerting solutions.
- Understanding of SRE principles, including SLIs, SLOs, availability, incident response, and root-cause analysis.
- Experience writing automation and operational tooling using Python and Bash.
- Strong understanding of networking fundamentals including TCP/IP, routing, DNS, TLS, and secure connectivity.
- Experience automating machine provisioning, configuration, deployment, and upgrades.
- Ability to independently troubleshoot complex issues spanning infrastructure, operating systems, networking, and application runtime.
- Strong ownership mindset and ability to take responsibility for systems from deployment through production operation.
Nice to Have
- Experience with Linux performance tuning including CPU pinning, NUMA, hugepages, IRQ affinity, and NIC tuning.
- Experience with high-performance or low-latency networking environments.
- Experience with DPDK or other user-space networking technologies.
- Experience operationally integrating FPGA or other hardware accelerators into Linux environments.
- Experience supporting software deployed in customer-managed or on-prem environments.
- Experience with PKI, certificate management, TLS/mTLS, and machine identity.
- Experience with Infrastructure as Code such as Terraform.
- Experience with GitHub Actions or similar CI/CD systems.
- Experience with vulnerability management and security/compliance initiatives such as SOC 2.
- Experience working in startup or high-growth engineering environments.
Why work with us?
- People-first management with minimal bureaucracy
- A friendly company culture, proven by employees who choose to return
- Flexible working hours
- 29 days of PTO (18 working days per year pluse all national holidays)
- 10 paid recovery days
- Full financial and legal support for independent contractors
- Free English classes, with native speakers or Ukrainian teachers
- Dedicated HR support
Our next steps
✅ Intro call with a Recruiter — ✅ Technical Interview — ✅ Interview with CTO and Head of Engineering — ✅ HR Interview — ✅ Offer