1 ساعت پیش | کد آگهی: 11850442
دستهبندی شغلی
موقعیت مکانی
تحصیلات
-
محل فعالیت
-
مزایا
-
مهارت ها و زبان ها
نوع همکاری
سایر اطلاعات
We are building and scaling a platform around AI agents digital services and multiple new lines of business As our systems grow reliability scalability observability and operational excellence become critical parts of the product itself We are looking for a Senior SRE DevOps Engineer who can help us design operate and continuously improve the infrastructure and engineering practices behind our services This role goes beyond maintaining clusters or deployment pipelines We are looking for someone who understands large scale software systems can identify the right operational and infrastructure patterns introduce better practices when needed and work closely with software engineers and external technical teams to make services reliable and production ready What You ll Do Design operate and continuously improve reliable and scalable production infrastructure Build and maintain Kubernetes based environments and containerized workloads Design and improve GitOps based deployment workflows using Argo CD and Helm Create and maintain reusable Helm charts and deployment standards across services Define and maintain SLIs SLOs SLAs error budgets and reliability targets for critical services Help define measurable technical and operational requirements for third party vendors and development partners Work with external teams to ensure their services meet agreed standards for availability latency monitoring scalability and incident response Design and maintain monitoring logging tracing dashboards and alerting systems Improve observability so engineering teams can quickly understand system health and diagnose production issues Troubleshoot complex production problems across applications infrastructure databases networking and distributed systems Perform root cause analysis and help implement long term fixes rather than relying on temporary operational workarounds Improve service resilience through proper timeout retry rate limiting failover and scaling strategies Perform capacity planning and help services scale as traffic and workloads increase Work closely with software engineers to make new services scalable observable resilient and production ready Consult development teams on architecture infrastructure databases caching messaging deployment strategies and operational tooling Help teams choose technologies and engineering practices based on actual technical and business requirements Define production readiness standards and review services before major releases Build reusable infrastructure and platform capabilities that simplify deployment and operation for development teams Automate repetitive operational processes and reduce manual intervention Improve backup recovery failover and disaster recovery processes where required Job Requirements Core Requirements Strong professional experience in Site Reliability Engineering DevOps Platform Engineering or production infrastructure preferably in large scale environments Strong hands on experience with Kubernetes and containerized production environments Strong experience with Helm including designing and maintaining reusable Helm charts Strong hands on experience with Argo CD and GitOps based deployment practices Good understanding of Kubernetes concepts including deployments and workloads services and ingress networking storage resource management autoscaling high availability access control and security Strong understanding of distributed systems scalability availability fault tolerance and production architecture Experience designing and operating high availability production services Strong understanding of GitOps and declarative infrastructure and deployment practices Experience designing deployment strategies such as rolling updates canary releases and controlled rollouts Strong experience with observability technologies such as Prometheus Grafana Elasticsearch OpenSearch ELK EFK Loki OpenTelemetry and distributed tracing platforms Strong understanding of metrics logs traces dashboards and actionable alerting Practical experience defining SLIs SLOs SLAs error budgets availability targets and reliability requirements Preferred Qualifications Experience in one or more of the following is considered a plus Large scale or high traffic platforms Multi cluster Kubernetes environments Advanced Helm deployment patterns Argo CD at scale Kubernetes operators and controllers Service mesh technologies API gateways Kafka and event driven architectures PostgreSQL or other large production databases Redis clusters Elasticsearch OpenSearch Ceph MinIO or other distributed storage platforms Secrets management solutions OpenTelemetry and distributed tracing تهران تهران عباس آباد بهشتی تمام وقت فناوری اطلاعات و ارتباطات مُهَیمن کارشناس ارشد جنسیت تفاوتی ندارد فناوری اطلاعات نرم افزار و سخت افزار
جویا کار این آگهی را از سایت
جابویژن
استخراج نموده است و هیچ مسئولیتی در قبال این آگهی ندارد.
دقت نمایید که کارفرما حق دریافت هیچ گونه وجهی از کارجو را نداشته و این امر خلاف قانون است. در صورت مشاهده این موارد یا سایر تخلفات با کلیک روی (گزارش آگهی) ما را در ارائه خدمات بهتر یاری نمایید.
در غیر این صورت میتوانید با کلیک بر روی دکمه "درج نظر" نظر خود را در مورد این آگهی ثبت کنید.
جهت اشتراک در شبکه های اجتماعی روی کلیدهای زیر کلیک کنید
همچنین میتوانید لینک کوتاه زیر را جهت دسترسی به صفحه فوق برای اشتراک گذاری کپی کنید
کپی کردن لینک
نظرات در مورد این آگهی: درج نظر