Job Description
Job DescriptionAbout the Role
This is a senior applied AI engineering role focused on building the AI systems layer behind a B2B process intelligence platform. You'll own the production systems that transform messy enterprise data into structured, reliable insights — sitting at the intersection of retrieval, orchestration, evaluation, and model integration. The work directly shapes the quality and usefulness of AI outputs delivered to real customers.
What You'll Do
-
Build and own production systems on top of hosted frontier LLM APIs (OpenAI, Anthropic, Gemini, Cohere).
-
Design and implement retrieval and context construction pipelines, including RAG, vector search, hybrid search, chunking, and reranking.
-
Create evaluation datasets, quality gates, regression tests, and LLM-as-judge workflows to measure and prevent model quality degradation.
-
Own structured extraction, model/provider selection, prompt optimization, and schema design to improve output quality.
-
Build and iterate on agent and tool orchestration workflows, including multi-step pipelines and planner/executor patterns.
-
Instrument systems for observability, cost tracking, latency monitoring, and failure mode handling.
-
Partner closely with backend, product, and forward-deployed engineering teams to solve real customer workflow problems.
What We're Looking For
-
3–6+ years of experience building and shipping production software, ML systems, or applied AI systems (not internal tooling or notebooks only).
-
Strong Python skills for building reliable services, data pipelines, eval workflows, and AI tooling.
-
Hands-on experience with LLM APIs, including prompting, structured outputs, embeddings, retrieval, and workflow orchestration.
-
Experience building evals, benchmarks, or CI/CD pipelines for model quality assessment.
-
Background in RAG, retrieval systems, or agent orchestration.
-
Breadth across multiple ML domains — not a narrow specialist in a single area.
-
Strong systems thinking: ability to reason end-to-end from raw data through model outputs to user-facing behavior.
-
Comfort with ambiguity — translating fuzzy product goals into experiments, implementations, and shipped improvements.
-
BS or higher in Computer Science, Engineering, or a related technical field.
-
Nice to have: experience with multimodal inputs (documents, screenshots, transcripts); familiarity with Go; experience with GCP, PostgreSQL, Redis, or Terraform.
Compensation & Benefits
Salary range: $180,000 – $275,000 USD annually. Visa sponsorship is not available for this role.
Location
On-site in New York, NY, United States. Candidates based in San Francisco may also be considered.