Assail is a Las Vegas–based cybersecurity startup building autonomous AI agents that deliver continuous, machine-speed penetration testing across APIs, web, and mobile applications.
<h2 class="text-text-100 mt-3 -mb-1 text-[1.375rem] font-bold">Senior AI Engineer, Ares Platform</h2> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]"><strong>Team:</strong> Ares AI Engineering <strong>Reports to:</strong> Ilir Osmanaj, VP of AI Engineering <strong>Location:</strong> Boston, MA (hybrid) or remote with overlap to ET working hours</p> <h3 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">Position summary</h3> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The Senior AI Engineer is a core builder on the team responsible for the agents and models that power Ares — Assail's autonomous offensive security platform for APIs, web applications, and mobile applications. This role works directly on Ares' named-agent architecture (Polemos, Hermes, Enyo, Momos, Dolos, Themis, Aletheia, Argus, Kratos), the model powering Ares, and the Javelin co-evolutionary self-training loop. The engineer will ship capabilities that move the platform forward across exploit chaining, multimodal vision, mobile coverage, self-improvement, and customer-facing accuracy.</p> <h3 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">Core tasks</h3> <ul class="[li_&]:mb-0 [li_&]:mt-1 [li_&]:gap-1 [&:not(:last-child)_ul]:pb-1 [&:not(:last-child)_ol]:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3"> <li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Agent development.</strong> Design, implement, and continuously improve the behavior and prompting of Ares' named agents, including orchestration patterns, hand-offs, planning loops, tool use, and shared memory.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Model training and fine-tuning.</strong> Contribute to the model powering Ares across data curation, SFT, preference optimization (DPO/GRPO-style), and evaluation. Own pieces of the training pipeline from dataset construction through eval.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Javelin loop.</strong> Extend the co-evolutionary self-training system that lets Ares learn from its own engagements and improve over time.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Self-improvement systems (ARES-420 and successors).</strong> Build false-positive detection, tiered skill learning (suppression rules, agent directives, code-patch proposals), and the infrastructure that routes proposed changes through human approval and back into the platform.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Evals.</strong> Design rigorous, security-specific evaluations covering OWASP Top 10 coverage, exploit chaining, finding accuracy, and agent reliability. Track performance over every model and agent change.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Multimodal and platform expansion.</strong> Contribute to vision capabilities, mobile (iOS/Android) coverage, and BYOK support shipping in Sidewinder and beyond.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Production reliability.</strong> Own latency, cost, observability, and failure-mode analysis for agents running in customer engagements. Partner with the platform team on Kubernetes-based deployment.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Customer-facing accuracy.</strong> Contribute to the live accuracy gauge and other surfaces where model and agent quality is exposed to customers.</li> </ul> <h3 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">Must-have skills</h3> <ul class="[li_&]:mb-0 [li_&]:mt-1 [li_&]:gap-1 [&:not(:last-child)_ul]:pb-1 [&:not(:last-child)_ol]:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3"> <li class="font-claude-response-body whitespace-normal break-words pl-2">5+ years building production ML/AI systems, with at least 2 years working directly on LLMs or LLM-powered agents.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Deep Python; strong, production-grade engineering practices (testing, code review, observability).</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Hands-on fine-tuning experience: SFT, preference optimization (DPO, GRPO, RLHF/RLAIF), data curation, and synthetic data generation.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Strong grasp of transformer architectures and the modern training stack (PyTorch, Hugging Face, DeepSpeed or FSDP, accelerate).</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Experience designing and shipping multi-agent or tool-using LLM systems in production — not just demos.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Rigorous eval design: building harnesses, tracking experiments, and making model/agent decisions based on data rather than vibes.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Inference optimization experience: vLLM or TensorRT-LLM, quantization, throughput/latency tradeoffs.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Comfort with retrieval pipelines, vector stores, and structured memory for agents.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Kubernetes and containerized deployment fluency.</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Genuine interest in offensive security and the ability to ramp quickly on OWASP Top 10, API security, web app pentesting, and mobile pentesting concepts. Direct offensive security background is a strong plus but not required.</li> </ul> <h3 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">Nice to have</h3> <ul class="[li_&]:mb-0 [li_&]