Senior/Staff Platform Observability Engineer
Employer · Espoo
Senior/Staff Platform Observability Engineer at Employer, based in Espoo. This is a permanent role with hybrid working.
- Salary
- Competitive
- Location
- Espoo · Hybrid
- Contract
- Permanent
- Posted
- 20 hours ago
- Closes
- 12 Sep 2026
- Sector
- Security
Reference j_76e357ec
About the role
ROLE HIGHLIGHTS: - Senior/Staff Platform Observability Engineer - Location: Espoo, Finland - Department: Platform & Release Engineering - Employment type: Permanent - Workplace model: Hybrid - Employment is subject to applicable security screening (incl. SUPO,) WHY THIS ROLE MATTERS: As a Platform Observability Engineer, you’ll make the health of ICEYE’s systems visible, actionable, and reliable. You’ll build and evolve the observability foundation across metrics, logs, and traces, helping teams detect issues early, understand failures faster, and operate their services with confidence. Your work will go beyond dashboards and alerts. You’ll reduce operational noise, eliminate blind spots, and give product teams the tools and autonomy to run reliable services. WHO WE ARE ICEYE is the world leader in sovereign intelligence from space. We deliver persistent monitoring capabilities to detect and respond to changes in any location on Earth. ICEYE owns the world's largest and most advanced SAR (synthetic aperture radar) satellite constellation. To our customers we provide intelligence with unmatched quality, latency and revisit times, in any weather, day or night. To governments who choose to operate their own constellation we provide this proven capability as a sovereign system. ICEYE-built constellations serve customers in defence and intelligence, environmental monitoring, insurance and emergency management. We enable fast decisions that contribute to a safer future. Founded and headquartered in Finland, ICEYE operates globally with over 1000 employees across Europe, North America, the Middle East, and Asia-Pacific. YOUR DAY-TO-DAY RESPONSIBILITIES - Own and continuously improve a unified, production-grade observability stack covering metrics, logs, and traces, giving every team a consistent, self-service way to understand and operate the health of their services. - Build and maintain trusted alerting by tuning alerts against clear SLOs/SLIs, reducing noise, improving ownership and ensuring real issues surface early. - Define, document, and drive adoption of instrumentation standards covering metric naming, label cardinality, structured logging, and distributed tracing with OpenTelemetry. - Enable product teams to diagnose and resolve incidents faster through well-designed dashboards, runbooks, hands-on support, and practical observability guidance. - Run the observability platform like a product, owning its roadmap, interfaces, reliability, scalability, and cost across areas such as retention, sampling, and cardinality. - Provide clear technical direction for observability across the engineering organisation, making pragmatic architectural trade-offs and helping teams adopt platform capabilities effectively. WHAT WE’RE LOOKING FOR Must haves: 1. Observability platform ownership at scale 6+ years in Platform Engineering, Observability, or infrastructure roles, with end-to-end ownership of a production observability stack (for example the LGTM stack - Loki, Grafana, Tempo, Mimir/Prometheus - or an equivalent metrics/logs/traces platform) at scale. - Instrumentation & telemetry: OpenTelemetry adoption, metric and label design, structured logging, and distributed tracing across complex, multi-team systems. - SLO/SLI & error budgets: dashboards and alerting that surface real problems early while minimizing noise and false positives. - Platform as a product: treating other engineering teams as customers, with clear interfaces and a roadmap shaped by their needs. - Alerting & reliability attitude: comfortable with on-call, driving blameless learning from incidents, and continuously tuning signal quality instead of letting alert fatigue set in. 2. Software engineering foundation - A background that includes time spent writing and shipping production software, bringing engineering instincts to building tooling and automation rather than only operating what already exists. - Proficiency in at least one backend language (for example Python or Go) for automation, tooling, and building internal observability capabilities. 3. Long-term ownership of your own decisions Demonstrated experience owning the long-term consequences of your own architectural and tooling decisions - having lived with what you built through its maintenance, upgrades, and failure modes, not just its initial rollout. 4. Communication Excellent written and verbal communication skills in English, able to explain complex system behavior clearly to both engineers and non-technical stakeholders. Nice to haves: 1. Platform & telemetry breadth - Rancher, Istio, and Kubernetes-native observability patterns (service mesh telemetry, sidecars, eBPF-based tooling). - Bridging hybrid environments across AWS and on-prem platforms such as vSphere/VxRail. - Managing cost and cardinality on large-scale telemetry pipelines (retention policies, sampling, downsampling). 2. Practices that spread beyond your own team - Defining SLO/SLI frameworks or observability standards…
Reference: j_76e357ec · Posted 20 hours ago · Closes 12 Sep 2026 · Listed via Employer
Apply for this job
This role is listed via Employer. Applications are handled on the employer's site.
Apply on employer siteOpens the employer's website in a new tab.
Safe applying: a genuine employer will never ask you to pay for a DBS check, training or equipment, or move you onto WhatsApp before you are hired. If this listing does, report it and do not pay anything.