senior service reliability engineer

3 ساعت پیش | کد آگهی: 11752805

دسته‌بندی شغلی

کارشناس تجارت الکترونیک

جنسیت و تاهل

خانم و آقا | مجرد و متاهل

موقعیت مکانی

تهران

تحصیلات

-

محل فعالیت

-

مزایا

-

نوع همکاری

تمام وقت

سایر اطلاعات

توضیحات آگهی

Key Responsibilities Service Reliability Engineering Design and implement strategies to improve service reliability availability resilience and operational stability Monitor and analyze service health across distributed applications APIs infrastructure and supporting components Define and track SLIs SLOs SLAs and error budgets for critical services Identify reliability risks single points of failure recurring incidents and systemic weaknesses Drive initiatives to reduce service degradation incident frequency and recovery time Contribute to resilience fault tolerance and operational readiness improvements Establish reliability standards and best practices across services and platforms Observability Service Visibility Design and enhance observability capabilities across metrics logs traces and events Develop centralized monitoring and telemetry strategies for critical services Build and maintain service health dashboards and reliability views Design effective alerting strategies with appropriate thresholds prioritization and noise reduction Improve visibility across distributed services APIs infrastructure and integrations Correlate telemetry and operational events to identify service impacting conditions Incident Analysis Reliability Improvement Perform advanced analysis of service incidents alerts and operational events Support incident response by providing deep service health and telemetry insights Lead or contribute to Root Cause Analysis RCA and post incident reviews Identify recurring failure patterns and reliability gaps Translate incident findings into concrete monitoring architecture and operational improvements Track reliability improvement actions through implementation and effectiveness Performance Capacity Engineering Analyze service performance latency throughput and resource utilization trends Identify performance bottlenecks and potential service degradation risks Support capacity planning and forecasting using operational and monitoring data Analyze system behavior under normal and abnormal operating conditions Provide recommendations for performance scalability and resource optimization Develop service reliability performance and capacity reports Service Health Operational Readiness Define and maintain service health indicators for critical business and technical services Establish operational readiness criteria for new and existing services Assess service dependencies and their potential impact on reliability Support reliability assessments for major changes releases and new service deployments Ensure monitoring alerting recovery and operational requirements are properly defined Collaborate with technical teams to improve overall service resilience and operational maturity Required Technical Skills Strong hands on experience in Service Reliability Production Operations or Site Reliability Engineering Strong experience with monitoring and observability platforms such as Prometheus Zabbix Grafana Splunk and Kibana Experience with centralized logging and telemetry platforms such as ELK Stack or Splunk Strong understanding of observability concepts metrics logs traces events and telemetry Solid understanding of distributed systems and service dependencies Strong Linux administration and performance troubleshooting skills Good understanding of network and infrastructure monitoring fundamentals Experience with application API and service level monitoring Experience with event management incident management and ticketing systems Understanding of service availability reliability and operational monitoring practices Reliability Engineering Knowledge Strong understanding of Service Reliability Engineering principles Practical knowledge of SLO SLI SLA and error budget frameworks Strong understanding of MTTR MTBF availability resilience and service continuity Experience with incident analysis RCA and post incident improvement Knowledge of alert optimization and signal to noise improvement Ability to identify systemic reliability risks and recurring failure patterns Understanding of fault tolerance redundancy graceful degradation and resilience Strong analytical and problem solving skills Understanding of service dependency and failure domain analysis Nice to Have Experience with APM and distributed tracing platforms Experience with OpenTelemetry Experience with automation and scripting using Python or Bash Exposure to cloud native and containerized environments Familiarity with Kubernetes and microservices architectures Experience with performance testing and capacity analysis Familiarity with ITIL ITSM processes Experience with service management or enterprise production environments Soft Skills Strong analytical and systems thinking mindset Calm and effective under incident pressure Cross team collaboration skills Clear reporting and documentation Work Conditions Participation in incident scenarios and reliability reviews Close collaboration with Service Operations teams Participation in incident reviews RCA and reliability improvement initiatives تهران تهران ونک تمام وقت همراه کسب و کارهای هوشمند کارشناس ارشد جنسیت تفاوتی ندارد اینترنت تجارت الکترونیک خدمات آنلاین

هشدار

جویا کار این آگهی را از سایت جاب‌ویژن استخراج نموده است و هیچ مسئولیتی در قبال این آگهی ندارد.
دقت نمایید که کارفرما حق دریافت هیچ گونه وجهی از کارجو را نداشته و این امر خلاف قانون است. در صورت مشاهده این موارد یا سایر تخلفات با کلیک روی (گزارش آگهی) ما را در ارائه خدمات بهتر یاری نمایید.
در غیر این صورت میتوانید با کلیک بر روی دکمه "درج نظر" نظر خود را در مورد این آگهی ثبت کنید.

نظرات در مورد این آگهی:

هیچ نظری برای این آگهی ثبت نشده است.