About the Project & Our Company
We are one of the biggest accounting companies in the world with 25,000 staff in China. We are executing a high-visibility, strategic project to implement a separate, dedicated cloud-based instance of our global accounting system within China. This project is crucial four future. The new system must utilize exclusively Chinese AI technologies Large Language Models (LLMs) (e.g., Baidu ERNIE, Alibaba Tongyi Qianwen) to adhere to regulatory requirements. This role joins a supportive global team of over 200 technology professionals.
Role Overview
We are seeking a proactive Site Reliability Engineer (SRE) to apply a specialized engineering approach to keeping our systems running. This role bridges the gap between the Development teams (who build the accounting system) Operations teams (who maintain it). You will ensure the system is reliable, scalable, efficient while operating on local Chinese cloud infrastructure. We are looking findividuals with a background in software engineering systems, prioritizing candidates with a strong academic background technical aptitude, including those with low experience (0-3 years).
Key Responsibilities
• Define Maintain Service Reliability: Establish monitkey Service Level Indicators (SLIs), set robust Service Level Objectives (SLOs), manage the ErrBudget to balance feature velocity against system stability.
• AI Reliability Monitoring: Monitthe operational health performance of integrated Chinese LLMs, specifically focusing on Inference Latency to prevent the AI from slowing down the core accounting software.
• Local Infrastructure Operations: Manage the operational requirements of the China instance, specifically overseeing the connection between global code local Chinese cloud providers (e.g., 21Vianet, Alibaba Cloud Huawei Cloud).
• Automation Toil Reduction: Proactively identify manual operations ("Toil") write code, scripts, automation tools to eliminate them, improving the systems "Self-Healing" capabilities.
• Engineering Excellence: Drive the adoption of modern DevOps practices, including automated testing, CI/CD practices, infrastructure as code, ensuring reliable scalable delivery.
• Compliance Integration: Work with security development teams to ensure all infrastructure configurations operational procedures comply with Chinas data security laws internal governance policies.
• Collaboration Communication: Document technical specifications clearly in both English Chinese, participate in agile ceremonies with the global team, requiring proactive English communication skills.
Qualifications & Experience
Essential Requirements:
• Bachelor’s Master’s degree in Computer Science, Software Engineering, Information Technology, a related STEM field from a leading university.
• 0-3 years of relevant work experience; strong academic record demonstrated technical aptitude are required.
• Proficiency in programming, particularly Python.
• Familiarity with systems administration, network fundamentals, cloud environments (such as AWS, Azure, GCP, local Chinese cloud services).
• Understanding of the software development lifecycle methodologies.
• Strong verbal written English communication skills feffective collaboration in a multinational team environment.
Highly Desirable (Bonus):
• Familiarity with DevOps practices tools, including CI/CD pipelines.
• Experience with containerization technologies (e.g., Docker, Kubernetes).
• Knowledge of MLOps practices relevant to the Chinese ecosystem.
• Understanding of the Chinese AI technology landscape local data center operations.
--------------------------------------------------------------------------------
SRE vs. Traditional Support:
The SRE role is an engineering function, meaning the focus is proactive, building systems so they do not break. In contrast, traditional IT support is reactive, primarily fixing issues after they happen, relying on manual checklists tickets. The SRE uses code to detect server hangs restart the system automatically, maximizing the systems ability to heal itself.
More