QA Analyst (AI Systems)
at Coderoad
- Seniority
- Staff Principal
- Location
- Latin America
- Posted
- 5d ago
at Coderoad
<h3></h3> <h2 data-path-to-node="4"><strong data-path-to-node="4" data-index-in-node="0">About CodeRoad</strong></h2> <p data-path-to-node="5">CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. From staff augmentation to dedicated IT teams and general software engineering, our nearshore technology services empower businesses to thrive in an ever-evolving digital landscape.</p> <h2 data-path-to-node="7"><strong data-path-to-node="7" data-index-in-node="0">About the Role</strong></h2> <p data-path-to-node="8">We are looking for a <strong data-path-to-node="8" data-index-in-node="21">Junior AI Agent Quality Engineer</strong> to help us build and validate the next generation of autonomous systems. In this role, you won't just be checking for bugs; you'll be evaluating the "brain" of our AI.</p> <p data-path-to-node="9">You will focus on the reliability of <strong data-path-to-node="9" data-index-in-node="37">Agentic AI</strong> and <strong data-path-to-node="9" data-index-in-node="52">RAG (Retrieval-Augmented Generation)</strong>. You’ll be responsible for ensuring our agents can plan tasks, use tools accurately, and recover from errors without "hallucinating." This is a perfect role for someone with a QA mindset who is passionate about Python and the future of LLM-based orchestration.</p> <h2 data-path-to-node="11"><strong data-path-to-node="11" data-index-in-node="0">Key Responsibilities</strong></h2> <ul data-path-to-node="12"> <li> <p data-path-to-node="12,0,0"><strong data-path-to-node="12,0,0" data-index-in-node="0">Agent Evaluation:</strong> Create question templates and Python scripts to test how well AI Agents and RAG instances solve complex tasks.</p> </li> <li> <p data-path-to-node="12,1,0"><strong data-path-to-node="12,1,0" data-index-in-node="0">KPI Tracking:</strong> Evaluate Agent performance using specific metrics: <strong data-path-to-node="12,1,0" data-index-in-node="65">Success Rate, Tool Use Accuracy, Planning Quality, and Autonomy</strong>.</p> </li> <li> <p data-path-to-node="12,2,0"><strong data-path-to-node="12,2,0" data-index-in-node="0">Dataset Creation:</strong> Collaborate with AI Engineers to build "Ground Truth" datasets and <strong data-path-to-node="12,2,0" data-index-in-node="85">Agentic Task Datasets</strong> to benchmark model improvements.</p> </li> <li> <p data-path-to-node="12,3,0"><strong data-path-to-node="12,3,0" data-index-in-node="0">API & Logic Testing:</strong> Write unit and integration tests for Python-based RESTful APIs and agent endpoints.</p> </li> <li> <p data-path-to-node="12,4,0"><strong data-path-to-node="12,4,0" data-index-in-node="0">Performance Testing:</strong> Run load and bulk testing (using <strong data-path-to-node="12,4,0" data-index-in-node="54">Locust or JMeter</strong>) to see how our AI handles high-volume requests.</p> </li> <li> <p data-path-to-node="12,5,0"><strong data-path-to-node="12,5,0" data-index-in-node="0">Input/Output Validation:</strong> Perform rigorous testing of agent payloads to catch prompt injection risks and logic failures.</p> </li> </ul> <hr data-path-to-node="13"> <h2 data-path-to-node="14"><strong data-path-to-node="14" data-index-in-node="0">Requirements</strong></h2> <ul data-path-to-node="15"> <li> <p data-path-to-node="15,0,0"><strong data-path-to-node="15,0,0" data-index-in-node="0">Agentic AI Focus:</strong> Basic understanding of <strong data-path-to-node="15,0,0" data-index-in-node="41">Agentic workflows</strong>, prompt engineering (ReAct prompts), and LLM orchestration.</p> </li> <li> <p data-path-to-node="15,1,0"><strong data-path-to-node="15,1,0" data-index-in-node="0">AI Observability:</strong> Familiarity with (or a strong desire to learn) tools like <strong data-path-to-node="15,1,0" data-index-in-node="76">LangSmith, Langfuse, or OpenTelemetry</strong>.</p> </li> <li> <p data-path-to-node="15,2,0"><strong data-path-to-node="15,2,0" data-index-in-node="0">Python Skills:</strong> Proficiency in <strong data-path-to-node="15,2,0" data-index-in-node="30">Python 3.10+</strong> for automation and data manipulation.</p> </li> <li> <p data-path-to-node="15,3,0"><strong data-path-to-node="15,3,0" data-index-in-node="0">API Fundamentals:</strong> Strong experience testing and validating RESTful APIs and JSON structures.</p> </li> <li> <p data-path-to-node="15,4,0"><strong data-path-to-node="15,4,0" data-index-in-node="0">Analytical Thinking:</strong> A "breaker" mindset—the ability to find edge cases where an AI might fail to follow instructions.</p> </li> <li> <p data-path-to-node="15,5,0"><strong data-path-to-node="15,5,0" data-index-in-node="0">Language:</strong> Advanced English (B2/C1) for global team collaboration and technical documentation.</p> </li> </ul> <hr data-path-to-node="16"> <h2 data-path-to-node="17"><strong data-path-to-node="17" data-index-in-node="0">Nice to Have</strong></h2> <ul data-path-to-node="18"> <li> <p data-path-to-node="18,0,0"><strong data-path-to-node="18,0,0" data-index-in-node="0">QA Experience:</strong> 1–3 years of experience in Software QA (Manual or Automation).</p> </li> <li> <p data-path-to-node="18,1,0"><strong data-path-to-node="18,1,0" data-index-in-node="0">Automation Tooling:</strong> Exposure to <strong data-path-to-node="18,1,0" data-index-in-node="32">Pytest</strong>, Playwright, or Selenium.</p> </li> <li> <p data-path-to-node="18,2,0"><strong data-path-to-node="18,2,0" data-index-in-node="0">AI Frameworks:</strong> Exposure to <strong data-path-to-node="18,2,0" data-index-in-node="27">LangChain, LangGraph, or Pydantic AI</strong>.</p> </li> <li> <p data-path-to-node="18,3,0"><strong data-path-to-node="18,3,0" data-index-in-node="0">Infrastructure:</strong> Basic knowledge of <strong data-path-to-node="18,3,0" data-index-in-node="35">Docker</strong> and Vector Databases.</p> </li> <li> <p data-path-to-node="18,4,0"><strong data-path-to-node="18,4,0" data-index-in-node="0">