Senior Engineer AI
Delta Exchange
Job Description
<h3>Role Summary</h3><p>We are looking for a Senior FullStack Engineer, AI to own and evolve <a href="https://himalayas.app/companies/delta-exchange">Delta Exchange</a>'s suite of AI-powered products. You will work across multiple production AI applications: conversational support agents, RAG-powered search, code-generation copilots, and trading-strategy assistants. This is a hands-on role where you architect solutions around our existing stack, build retrieval pipelines, improve frontends with product sense, optimise inference costs, and run evaluations, collaborating closely with the product team.</p><h3>Key Responsibilities</h3><ul>
<li>Design, build, and maintain production AI applications end-to-end: backend, frontend, and inference services.</li>
<li>Architect RAG systems using vector databases, embedding models, and chunking strategies optimised for accuracy and latency.</li>
<li>Build agentic workflows with tool/function calling, multi-step reasoning, and structured output parsing, with accuracy and control as priority.</li>
<li>Write and iterate on system prompts, few-shot examples, and prompt chains to maximise output quality.</li>
<li>Implement function calling, tool-use patterns, and structured JSON/XML output handling using frontier and lightweight models from providers like Anthropic and OpenAI.</li>
<li>Drive cost optimisation: model selection, caching, token budgeting, and request batching at scale.</li>
<li>Build and maintain evaluation frameworks to measure accuracy, relevance, hallucination rates, and regression across prompt and model changes. Experience with observability tools (Sentry, Opik, etc.) is a must.</li>
<li>Work with message queues (RabbitMQ), caching layers (Redis), and relational databases (PostgreSQL) powering AI service backends.</li>
<li>Deploy and manage AI services on Kubernetes with CI/CD pipelines on AWS/GCP.</li>
<li>Integrate AI capabilities with third-party platforms (Telegram bots, chat widgets, etc.).</li>
<li>Contribute to architectural decisions: model selection, hosting (cloud APIs vs. self-hosted), and build-vs-buy trade-offs.</li>
</ul><h3>Required Skills & Experience</h3><ul>
<li>5+ years shipping production software systems.</li>
<li>2 years building AI/LLM-powered applications end-to-end with real users and volume. Not prototypes.</li>
<li>Strong experience with RAG architectures: vector databases, embedding models, chunking/indexing strategies, and retrieval evaluation.</li>
<li>Deep understanding of LLM capabilities and limitations: prompt engineering, function/tool calling, structured outputs, context window management, and multi-turn conversations.</li>
<li>Experience with LLM provider APIs and abstraction layers (OpenAI, Anthropic, LiteLLM, OpenRouter, or similar).</li>
<li>Proficiency in Python (Flask/FastAPI) and/or Node.js/TypeScript (Next.js, Vercel AI SDK). Golang experience is a plus.</li>
<li>Hands-on experience building evals, tracking quality metrics, and debugging non-determ