Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this AI Engineer role.
Rezi rewrites your resume against NetBrain's job description. Free.

Tailor your resume to this AI Engineer role.
Rezi rewrites your resume against NetBrain's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the AI Engineer posting at NetBrain — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the AI Engineer posting at NetBrain — free, in seconds.
About the Role
We are seeking a Senior AI Engineer to design and build production-grade agent and RAG systems that power intelligent, reliable automation across our platform. This role requires hands-on engineering, system-level thinking, and ownership from architecture to production reliability. The ideal candidate thrives in ambiguity, moves quickly from prototype to production, and prioritizes quality, safety, and real-world impact.
Responsibilities
- Design and implement core capabilities for an enterprise-grade Agent platform, including orchestration patterns, tool execution, context and memory management, and safety guardrails.
- Design enterprise-grade Agent execution and governance mechanisms, including Human-in-the-Loop approval workflows, multi-tenant permission isolation, policy enforcement, and secure execution controls.
- Build reusable Agent Skills, standardized tool interfaces, and a scalable tool ecosystem integrated with NetBrain platform capabilities.
- Design and implement LLM post-training strategies, including domain-specific Supervised Fine-Tuning (SFT), DPO/RLHF-based preference alignment, and parameter-efficient fine-tuning techniques like LoRA.
- Build an Agent self-learning feedback loop to convert production execution traces, user feedback, and evaluation results into high-quality datasets for continuous improvement.
- Analyze and optimize LLM behavior across instruction following, tool calling, structured output generation, contextual understanding, reasoning stability, and hallucination mitigation.
- Build production-grade LLM and Agent evaluation frameworks and automated regression pipelines.
- Establish release quality gates and hallucination-detection mechanisms for AI features.
- Build comprehensive AI system observability capabilities, including distributed tracing, structured logging, metrics, dashboards, and alerting.
- Rapidly diagnose and resolve production AI failures.
- Design and implement highly reliable backend services for production AI and Agent workloads.
- Continuously optimize latency, throughput, token consumption, and infrastructure cost.
- Independently diagnose and resolve complex AI system issues.
- Lead technical design for critical modules and system-level capabilities.
- Drive technical improvements based on production data, evaluation results, and benchmarks.
- Prototype, benchmark, and productionize emerging technologies such as GraphRAG, Knowledge Graphs, MCP, LLM Post-Training, and Agent Self-Learning.
- Continuously evaluate Agent frameworks and supporting infrastructure, providing technical recommendations for platform architecture evolution and product technology strategy.
- Stay current with developments in LLM and Agent technologies and rapidly translate promising technologies into production-ready engineering capabilities.
Requirements
- Bachelor's degree or higher in Computer Science, Artificial Intelligence, Electrical Engineering, or a related technical field. Master's or Ph.D. preferred; equivalent practical experience will also be considered.
- 3+ years of experience in software engineering, machine learning, or applied AI.
- 2+ years building, deploying, and operating production-grade LLM or Agent applications.
- Must have delivered at least one LLM-powered feature end-to-end and owned its ongoing operation and improvement after production launch.
- Deep understanding of Agent architectures and LLM behavioral characteristics.
- Hands-on experience building multi-step workflows involving reasoning, tool execution, state management, structured outputs, validation, and error recovery.
- Proven ability to diagnose and resolve production LLM/Agent failures.
- Strong Python and distributed backend engineering skills.
- Hands-on experience designing evaluation systems for LLM applications.
- Strong understanding of security risks associated with LLM and Agent applications.
- Ability to independently design, implement, debug, deploy, and operate complex production systems.
- Strong experience with RAG and advanced retrieval systems.
- Familiarity with Agent frameworks such as LangGraph, LangChain, AutoGen, and LlamaIndex.
- Experience designing Agent runtime mechanisms, including Human-in-the-Loop workflows.
- Experience with LLM fine-tuning, including LoRA or other parameter-efficient fine-tuning techniques.
- Experience applying LLM technologies to networking, infrastructure, cybersecurity, observability, or other complex technical domains.
- Fluent in both English and Chinese, with strong cross-regional communication and collaboration skills.
Skills
- Agent architectures
- LLM behavioral characteristics
- Instruction following
- Tool-calling behavior
- Context sensitivity
- Multi-step workflows
- Reasoning
- Tool execution
- State management
- Structured outputs
- Validation
- Error recovery
- Production LLM/Agent failure diagnosis and resolution
- Hallucinations
- Incorrect tool calls
- Retrieval-quality degradation
- Agent loops
- Structured-output failures
- Latency regressions
- Prompt changes
- Model changes
- Python
- Distributed backend engineering
- API development
- Service development
- Asynchronous programming
- Concurrent programming
- Retries
- Timeouts
- Caching
- Rate limiting
- Testing
- Logging
- Cross-service performance debugging
- LLM application evaluation systems
- Dataset construction
- Metric definition
- Regression testing
- Release quality gates
- Security risks associated with LLM and Agent applications
- Prompt injection
- Data leakage
- Unsafe tool execution
- Permission boundaries
- Uncontrolled Agent autonomy
- Production system design, implementation, debugging, deployment, and operation
- RAG
- Advanced retrieval systems
- Embeddings
- Vector search
- Hybrid search
- Reranking
- Chunking strategies
- Grounding mechanisms
- Citation mechanisms
- Multi-hop retrieval
- Knowledge Graphs
- GraphRAG
- LangGraph
- LangChain
- AutoGen
- LlamaIndex
- MCP (Model Context Protocol)
- Agent runtime mechanisms
- Human-in-the-Loop workflows
- Risk classification
- Approval workflows
- Interruption workflows
- Pause/resume workflows
- Context management for long-running Agent workflows
- State persistence
- History compression
- Memory systems
- LangSmith
- LLM fine-tuning
- LoRA
- Parameter-efficient fine-tuning
- LLM technologies applied to networking
- LLM technologies applied to infrastructure
- LLM technologies applied to cybersecurity
- LLM technologies applied to observability
Location
- Remote
Work Type
- Full-time
Experience Level
- Senior
- 3+ years of experience in software engineering, machine learning, or applied AI
- 2+ years building, deploying, and operating production-grade LLM or Agent applications
Education Level
- Bachelor's degree or higher in Computer Science, Artificial Intelligence, Electrical Engineering, or a related technical field
- Master's or Ph.D. preferred
Salary/Compensations
- CAD $130,000 - CAD $165,000 + Bonus
Benefits
- RRSP
- Medical coverage
- Dental coverage
About the Company
- Founded in 2004, NetBrain is the leader in no-code network automation.
- Its ground-breaking Next-Gen platform provides IT operations teams with the ability to scale their hybrid multi-cloud connected networks by automating the processes associated with Diagnostic Troubleshooting, Outage Prevention and Protected Change Management.
- Today, over 2,500 of the world’s largest enterprises and managed services providers leverage NetBrain’s platform.
Equal Opportunity
- NetBrain invites all interested and qualified candidates to apply for employment opportunities.
- Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability, protected veteran status, or other characteristics protected by law.
- If you have a disability that prevents or limits your ability to use or access the site, or if you require any other accommodation in the application process due to a disability, you may request a reasonable accommodation.
- In compliance with applicable laws, NetBrain conducts holistic, individual background reviews in support of all hiring decisions.
- It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.