Senior Support Engineer - Dublin
Posted 4 hours 58 minutes ago by Doist
The Technical Support team ensures that developers and enterprises can reliably build mission critical solutions using OpenAI models. We provide technical guidance, resolve complex issues, and support customers in maximizing value from our highly capable models. Working closely with Technical Success, Product, Engineering and others, we deliver the best possible experience at scale and leverage AI to automate and scale support operations.
About the RoleWe are looking for a Senior Support Engineer to collaborate directly with strategic enterprise accounts and product teams, helping solve the most difficult problems faced by our customers. In this role you will be part of the top technical troubleshooting team at OpenAI, the last line of defense before the core Engineering team. The role is based in Dublin, Ireland and follows a hybrid work model of three days in the office each week with relocation assistance for new employees.
Key Responsibilities- Design and run operational processes to monitor top strategic customers and support a 24x7 response team.
- Work closely with Infrastructure and Engineering teams to deliver the best possible experience to customers at scale.
- Identify and implement opportunities to scale support operations by leveraging automation and advances in AI technology.
- Configure and use advanced monitoring and alerting workflows to proactively detect customer impacting issues in real time.
- Participate in reliability reviews and readiness for new features, launches, or strategic customer updates, ensuring monitoring, alerting and fallback plans are in place.
- Design and refine incident response processes and documentation across strategic customers, engineering and support teams.
- Analyze operational metrics and incident RCAs to identify areas for improvement and recommend enhancements to dashboards and support workflows.
- Provide support coverage during holidays and weekends based on business needs.
- Bachelor's degree in Computer Science or a related field.
- 5+ years of experience in technical operations roles such as SRE/NOC, designing monitoring systems and resolving production issues in fast paced, mission critical environments.
- Proven record of troubleshooting complex technical problems at the systems level.
- Deep familiarity with modern monitoring, alerting and observability practices, including metrics, logging, and tracing for distributed systems.
- Experience leading incident response for high severity outages, performing real time incident coordination and root cause analysis, and driving follow ups such as post mortems.
- Knowledge of industry best practices for incident management and fault diagnosis.
- Strong scripting or software engineering skills (e.g., Python) for automating repetitive tasks and integrating tools.
- Solid understanding of cloud infrastructure and distributed systems fundamentals; comfortable with cloud services, load balancers, databases and containerized applications.
- Effective cross functional collaboration in a high trust environment, with strong communication skills to explain technical issues to both engineering and non technical stakeholders.
- Annual salary and generous equity.
- Medical, dental and vision insurance for employee and family.
- Medical mental health and wellness support.
- PRSA plan with 8% employer matching.
- Unlimited time off.
- Annual learning & development stipend ($1,500USD equivalent per year).
OpenAI is an equal opportunity employer and does not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or any other protected characteristic. We provide reasonable accommodations to applicants with disabilities. Background checks will be administered in accordance with applicable law.