Evals / model evaluation jobs
As of Oct 8, 2026, 4,147 of the 411,827 job listings in the Velza AI Job Search Tool mention Evals / model evaluation (1.0%). 1,433 of them (35%) say the job is remote. Most of these listings are in the United States (2,106), India (328) and the United Kingdom (180).
Most recent listings
Sr Pricing Analyst
Assurant · United States Virtual · Posted Oct 7, 2026
- Databricks
- Evals / model evaluation
External Communications Manager
AstraZeneca · UK - London · Posted Oct 7, 2026
- Evals / model evaluation
Director, Oncology Commercial Data Science & AI Products
AstraZeneca · US - Gaithersburg - MD · Posted Oct 7, 2026
- AI in title
- AI agents / agentic AI
- AI automation / workflow automation
- AI governance / responsible AI
- AI literacy / AI fluency
- Amazon Bedrock
- ChatGPT / OpenAI
- Claude / Anthropic
- Databricks
- Evals / model evaluation
- Gemini
- Generative AI
- Human-in-the-loop
- LangChain / LangGraph
- Large Language Models (LLMs)
- LlamaIndex
- MLOps / LLMOps
- Multimodal AI
- Prompt engineering
- Retrieval-augmented generation (RAG)
- Snowflake
- Vector databases / embeddings
Senior Machine Learning Engineer
Autodesk · Toronto, ON, CAN · Posted Oct 7, 2026
- AI in title
- AI agents / agentic AI
- Evals / model evaluation
- MLOps / LLMOps
- Machine learning
- Snowflake
AI/ML Engineer (Computer Vision)
CACI · Ashburn, VA, US · Posted Oct 7, 2026
- AI in title
- Computer vision
- Deep learning
- Evals / model evaluation
- Fine-tuning
Lead AI/ML Engineer (Computer Vision)
CACI · Ashburn, VA, US | Remote (Any State) · Posted Oct 7, 2026
- AI in title
- Computer vision
- Deep learning
- Evals / model evaluation
- Fine-tuning
AI Engineer 5 (Gen AI Platform Services: Agentic AI)
Capital One · San Jose, CA | San Francisco, CA | McLean, VA | Cambridge, MA | New York, NY · Posted Oct 7, 2026
- AI in title
- AI agents / agentic AI
- AI governance / responsible AI
- Evals / model evaluation
- Generative AI
- Hugging Face
- Human-in-the-loop
- Large Language Models (LLMs)
- Machine learning
- Vector databases / embeddings
AI Engineer 5 (Gen AI Platform Services: Agentic Systems)
Capital One · New York, NY | San Francisco, CA | McLean, VA | Cambridge, MA | San Jose, CA · Posted Oct 7, 2026
- AI in title
- AI agents / agentic AI
- AI governance / responsible AI
- Evals / model evaluation
- Generative AI
- Hugging Face
- Human-in-the-loop
- Large Language Models (LLMs)
- Machine learning
- Vector databases / embeddings
AI Engineer 5 (Gen AI Platform Services: Agentic AI, Guardrails, Evaluations)
Capital One · New York, NY | San Francisco, CA | McLean, VA | Cambridge, MA | San Jose, CA · Posted Oct 7, 2026
- AI in title
- AI agents / agentic AI
- AI governance / responsible AI
- Evals / model evaluation
- Generative AI
- Hugging Face
- Human-in-the-loop
- Large Language Models (LLMs)
- Machine learning
- Vector databases / embeddings
AI/ML Engineer
Cardinal Health · US-Nationwide-FIELD · Posted Oct 7, 2026
- AI in title
- AI agents / agentic AI
- AI automation / workflow automation
- AI governance / responsible AI
- ChatGPT / OpenAI
- Claude / Anthropic
- Evals / model evaluation
- Gemini
- Generative AI
- LangChain / LangGraph
- Large Language Models (LLMs)
- MLOps / LLMOps
- Machine learning
- Prompt engineering
- Retrieval-augmented generation (RAG)
- Vector databases / embeddings
Agentic AI Engineer Lead
Carrier · CAI23: Carrier-Indianapolis, 7310 West Morris Street, Indianapolis, IN, 46231 USA | CAG10: ALC HQ, 1025 Cobb Place Boulevard, Kennesaw, GA, 30145 USA · Posted Oct 7, 2026
- AI in title
- AI agents / agentic AI
- AI governance / responsible AI
- Evals / model evaluation
- Fine-tuning
- Generative AI
- Large Language Models (LLMs)
- MLOps / LLMOps
- Machine learning
- Model Context Protocol (MCP)
- Prompt engineering
- Retrieval-augmented generation (RAG)
- Vector databases / embeddings
Machine Learning Engineer - CTO innovations
Cisco · San Francisco, California, US | San Jose, California, US · Posted Oct 7, 2026
- AI in title
- Evals / model evaluation
- Generative AI
- LangChain / LangGraph
- Large Language Models (LLMs)
- Machine learning
- Prompt engineering
- Retrieval-augmented generation (RAG)
- Snowflake
Software Engineer - (Media SRE)
Cisco · Stockholm, Sweden · Posted Oct 7, 2026
- AI agents / agentic AI
- AI governance / responsible AI
- Evals / model evaluation
- Human-in-the-loop
- Large Language Models (LLMs)
- Model Context Protocol (MCP)
- Prompt engineering
- Retrieval-augmented generation (RAG)
Product Director, DevOps & CMO Performance
GSK · UK - Hertfordshire - Stevenage · Posted Oct 7, 2026
- AI agents / agentic AI
- AI automation / workflow automation
- AI governance / responsible AI
- Evals / model evaluation
- MLOps / LLMOps
- Vector databases / embeddings
Senior Director, Analyst – Software Engineering for AI and Agentic Applications (Remote - Canada) / Directeur principal, analyste – Ingénierie logicielle pour l’IA et les applications agentiques (Télétravail – Canada)
Gartner · Remote - Canada · Posted Oct 7, 2026
- AI in title
- AI agents / agentic AI
- AI literacy / AI fluency
- Evals / model evaluation
- Large Language Models (LLMs)
Software Engineer III, Merchant Shopping, Intelligence and Agents
Google · Zürich, Switzerland · Posted Oct 7, 2026
- AI agents / agentic AI
- Evals / model evaluation
- Large Language Models (LLMs)
- Machine learning
- NLP
- Prompt engineering
Staff Software Engineer, Ads and Commerce Advisor Platform
Google · Mountain View, CA, USA · Posted Oct 7, 2026
- AI agents / agentic AI
- Computer vision
- Evals / model evaluation
- Fine-tuning
- Generative AI
- Large Language Models (LLMs)
- Machine learning
- Multimodal AI
- NLP
Software Engineer III, AI/ML YouTube Shopping Creator
Google · Zürich, Switzerland · Posted Oct 7, 2026
- AI in title
- Evals / model evaluation
- Large Language Models (LLMs)
- Machine learning
- NLP
Senior Business Data Scientist, Marketing Science, Measurement and Optimization
Google · New York, NY, USA · Posted Oct 7, 2026
- AI in title
- Evals / model evaluation
- Machine learning
Senior Software Engineer, Google Health
Google · Mountain View, CA, USA · Posted Oct 7, 2026
- AI agents / agentic AI
- Evals / model evaluation
- Generative AI
- NLP
Software Engineer III, AI/ML, Omni-channel Shopping Ads Quality
Google · Mountain View, CA, USA · Posted Oct 7, 2026
- AI in title
- Evals / model evaluation
- Generative AI
- Machine learning
- NLP
Software Engineering Manager, YouTube Ads Machine Learning
Google · Mountain View, CA, USA · Posted Oct 7, 2026
- AI in title
- Computer vision
- Evals / model evaluation
- Fine-tuning
- Generative AI
- Large Language Models (LLMs)
- Machine learning
- Multimodal AI
- NLP
Software Engineer, Content Safety
Google · Singapore · Posted Oct 7, 2026
- AI agents / agentic AI
- AI governance / responsible AI
- Computer vision
- Evals / model evaluation
- Generative AI
- Large Language Models (LLMs)
- Machine learning
- Multimodal AI
- NLP
- Vector databases / embeddings
Senior Software Engineering Manager, Google Cloud Storage Semantic Search
Google · Seattle, WA, USA | Sunnyvale, CA, USA · Posted Oct 7, 2026
- Evals / model evaluation
- Fine-tuning
- NLP
- Vector databases / embeddings
Software Engineer III, AI/ML, TPU Efficiency, YouTube
Google · Mountain View, CA, USA · Posted Oct 7, 2026
- AI in title
- Evals / model evaluation
- NLP
- Vector databases / embeddings
Evals / model evaluation jobs by country
- United States2,106
- India328
- United Kingdom180
- Canada147
- Singapore84
- Germany53
- Ireland48
- Poland44
- Spain41
- China38
- Philippines37
- South Korea35
- Brazil32
- France28
- Japan28
- Taiwan28
- Israel27
- Australia26
- Portugal26
- Netherlands25
- Hong Kong24
- Indonesia23
- Vietnam23
- Sweden22
- Mexico21
Browse more topics
- Jobs in North America
- Jobs in Latin America
- Jobs in Europe
- Jobs in the Middle East & Africa
- Jobs in Asia Pacific
- Remote jobs
- AI agents / agentic AI jobs
- Large Language Models (LLMs) jobs
- Machine learning jobs
- AI literacy / AI fluency jobs
- Generative AI jobs
- Claude / Anthropic jobs
- AI automation / workflow automation jobs
- Vector databases / embeddings jobs
- Microsoft Copilot jobs
- ChatGPT / OpenAI jobs
- Retrieval-augmented generation (RAG) jobs
- Snowflake jobs
How this list works: we read the public job boards that employers publish on Greenhouse, Lever, Ashby, SmartRecruiters and Workday, plus Google's and Microsoft's own career sites. Locations, regions and remote status come from what each posting says. AI skill labels are matched from each posting's text. Velza is not the employer and does not take applications. Always confirm details on the employer's site.