Senior System Engineer , Zalo
Role and Development Direction
Build and operate a shared AI platform that enables teams to use AI safely, reliably, with proper controls, and in a cost-efficient manner. There are currently approximately 800 licensed users across platforms such as ChatGPT, Gemini, and Claude. This role serves as the technical focal point for standardizing access, managing the license lifecycle, controlling costs, and mitigating operational risks;
In the initial phase, the focus will be on AI access management, automating control measures, monitoring operations, and supporting teams in their AI usage. As the number of AI applications deployed in production and the level of agent autonomy increase, the scope may expand to include agent/tool management, isolated execution environments, integration of evaluation results into CI/CD pipelines, context/RAG services, and internal model operations, based on approved business needs and business cases.
🤖 What you will do
- Build and operate a multi-provider AI gateway/platform, including quota, SSO/RBAC, audit logs, DLP/redaction and observability;
- Manage the lifecycle of providers, models, licenses, connectors, plugins and MCP servers, from technical evaluation, proposal and coordination with relevant stakeholders for approval and access provisioning, to periodic review, revocation or replacement;
- Receive AI use cases, classify risks and advise on suitable deployment models; coordinate with ZSEC, IT, HRBP and Procurement when necessary;
- Standardize AI tooling for developers and business teams through portals, APIs/SDKs, secure configurations, documentation, standard implementation templates and self-service processes;
- Establish and enforce requirements for coding agents to connect to models/providers through the AI Proxy in accordance with company policies;
- Monitor stability, usage, latency, error rates and costs; troubleshoot incidents, cost spikes, provider changes and signs of data leakage.
👾 What you will need
- Bachelor’s degree in Information Technology or a related field;
- At least 3 years of experience operating systems/services in production as a System Administrator, IT Infrastructure Engineer, DevOps Engineer or SRE;
- Proficient in Linux and basic networking; experienced with commonly used infrastructure components such as web servers/reverse proxies, containers (Docker/Podman), account and access management (SSO, LDAP/AD or IAM), and secret/credential management;
- Experience in monitoring and troubleshooting: reading logs, setting up basic dashboards and alerts, and monitoring service usage and costs;
- Proficient in scripting (Python, Bash or equivalent) to automate repetitive operational tasks; basic understanding of Git, CI/CD and software development processes to collaborate with developer teams;
- Experience using LLM APIs or AI tools (chatbots, coding assistants) at work and willingness to gain in-depth knowledge of AI gateways, RAG and agent workflows;
- Basic understanding of access control, data protection, audit logging and security incident response processes;
- Ability to collaborate with multiple teams and write clear operational documentation and user guides.