Socium - Teams Done Differently
Experts at deploying talented teams across technology enabled businesses.
شرح موقعیت
Location: Abu Dhabi (Onsite) Contract Duration: Initial 12 months extendable contract Key Responsibilities Design and implement AI/LLM evaluation frameworks, benchmarks, and test datasets. Develop metrics and automated testing for accuracy, relevance, groundedness, hallucination, safety, and reliability. Evaluate LLMs, RAG pipelines, agents, and multimodal AI applications. Build regression testing and quality gates for AI/ML releases. Analyze evaluation results and identify opportunities for model and system improvements. Develop scalable evaluation pipelines and services using cloud-native and MLOps practices. Collaborate with ML engineers, data scientists, software engineers, and product teams. 4–6 years of experience in AI/ML Engineering, Data Science, NLP, or a related field. Strong Python and software engineering skills. Hands-on experience with LLMs, Generative AI, RAG, embeddings, and vector databases. Experience with AI/LLM evaluation, benchmarking, model testing, or quality measurement. Familiarity with LLM-as-a-Judge, RAG evaluation, prompt testing, and human-in-the-loop evaluation. Experience with AWS and/or Azure, Docker, Kubernetes, MLflow, and CI/CD. Strong SQL skills and experience with APIs and modern AI/ML architectures.