
ADIB
LLM Research & Validation Specialist
🇦🇪 UAE, UAE🏢 On-site
LLMRAGPythonAI SafetyMachine LearningAgentic AIModel Validation
At a Glance
- Category
- 💻 Technology
- Level
- Mid-Level
- Type
- Full-time
75% Get Rejected
- See if your CV passes ADIB's ATS filters
- Get AI-rewritten bullet points
- Download Gulf-ready CV
60 seconds. $5.88 one-time.
Job title: LLM Research & Validation Specialist
Location: Abu Dhabi, UAE
Role purpose:
- Lead frontier research, quantitative evaluation and independent validation of large language models, multimodal models, retrieval-augmented generation systems and agentic AI used or proposed by ADIB. Translate mathematical and scientific methods into reproducible validation tests, challenger analyses, runtime controls and decision-useful evidence for model governance.
- The role combines deep technical research with second-line effective challenge.
- It is expected to build validation toolkits and evaluation harnesses, independently assess conceptual soundness and production behaviour, and communicate material limitations clearly to technical teams, senior management and governance forums.
- The role does not own model development or production approval.
Key accountabilities /responsibilities:
- Lead independent validation of LLM, multimodal, RAG and agentic AI use cases across design, implementation, deployment and ongoing monitoring.
- Assess transformer architecture, tokenisation, embeddings, attention, context-window behaviour, decoding, fine-tuning, alignment, quantisation and inference configuration.
- Design reproducible evaluation harnesses, golden datasets, adversarial suites, counterfactual tests, canary sets and statistically defensible acceptance criteria.
- Evaluate task performance, hallucination and factuality, calibration, robustness, stability, long-context behaviour, retrieval quality, grounding, citation faithfulness and uncertainty.
- Perform deep testing of prompt injection, indirect injection, data leakage, tool-use safety, excessive agency, multi-step failure propagation, kill-switches and human oversight.
- Apply probability, statistics, optimisation, information theory, numerical methods and experimental design to develop challenger tests and quantify uncertainty.
- Review data provenance, representativeness, contamination, benchmark validity, leakage, drift and limitations of synthetic or LLM-generated evaluation data.
- Build and maintain reusable Python-based validation tooling, automated test pipelines, experiment tracking, results repositories and technical documentation.
- Conduct structured research on emerging model architectures, interpretability, mechanistic analysis, scalable oversight, model evaluation and AI safety methods.
- Independently challenge model owners, vendors and developers, document findings, propose risk-based restrictions and track remediation without assuming first-line ownership.
- Prepare validation reports, research notes, standards, committee papers and senior-management briefings that clearly distinguish evidence, judgement and residual uncertainty.
- Mentor junior validators, improve team methodology and support knowledge transfer across Model Risk
Education and experience:
- Master's degree in Theoretical Physics, Applied Physics, Mathematics, Applied Mathematics or a closely related quantitative discipline is required. A PhD or research-intensive master's is strongly preferred.
- Typically, one to three years of relevant experience in AI research, machine learning, quantitative modelling, model validation, scientific computing or a closely related field. Exceptional research profiles may be considered based on demonstrated capability.
- Deep understanding of probability, statistics, linear algebra, optimisation, numerical computation, experimental design and uncertainty quantification.
- Strong understanding of transformers, LLM training and inference, embeddings, RAG, fine-tuning, alignment, evaluation, agentic systems and AI safety failure modes.
- Advanced Python proficiency and experience with scientific and ML libraries. Exposure to PyTorch, Hugging Face, evaluation frameworks, experiment tracking, SQL, Git and cloud AI platforms is expected.
- Ability to read research papers critically, reproduce methods, design-controlled experiments and convert findings into bank-grade validation evidence.
- Experience with red teaming, adversarial testing, interpretability, calibration, robustness, privacy, security or model risk management is strongly advantageous.
- Excellent technical writing and communication, including the ability to explain mathematical concepts, assumptions and limitations to non-specialist stakeholders.
- Banking experience is advantageous but not mandatory. The role requires willingness to develop knowledge of financial services, Islamic banking, CBUAE expectations and ADIB governance.
Indicative success measures:
- Validation conclusions are reproducible, evidence-based and proportionate to use-case risk.
- Reusable evaluation assets and automation measurably improve validation coverage, consistency and efficiency.
- Material LLM and agentic risks are identified early, clearly communicated and translated into actionable controls or use restrictions.
- Research outputs strengthen ADIB validation methodology and remain traceable to tested evidence rather than unsupported claims.
- Stakeholders receive constructive, independent challenges while second-line ownership and decision rights remain clear.
Requirements
- •Master's degree in Theoretical Physics, Applied Physics, Mathematics, Applied Mathematics or a closely related quantitative discipline
- •PhD or research experience in quantitative disciplines preferred
- •Deep understanding of transformer architecture, tokenization, embeddings, and attention
- •Proficiency in probability, statistics, optimization, and information theory
- •Ability to build and maintain reusable Python-based validation tooling
- •Experience with adversarial suites, counterfactual tests, and canary sets
- •Strong ability to communicate material limitations to senior management and governance forums
Responsibilities
- •Lead independent validation of LLM, multimodal, RAG, and agentic AI use cases
- •Assess transformer architecture, fine-tuning, alignment, quantization, and inference configuration
- •Design reproducible evaluation harnesses, golden datasets, and statistically defensible acceptance criteria
- •Evaluate task performance, hallucination, factuality, calibration, robustness, and retrieval quality
- •Perform deep testing of prompt injection, data leakage, tool-use safety, and multi-step failure propagation
- •Review data provenance, representativeness, contamination, and benchmark validity
- •Build and maintain automated test pipelines, experiment tracking, and results repositories
- •Conduct structured research on emerging model architectures, interpretability, and AI safety methods
Related Jobs4 similar jobs
Guides for this role
Apply Now
75% Get Rejected
- See if your CV passes ADIB's ATS filters
- Get AI-rewritten bullet points
- Download Gulf-ready CV
60 seconds. $5.88 one-time.
GCC Info
Company

ADIB
5000+ employees
ADIB is a leading Islamic bank offering a wide range of banking and financial services. It serves individuals, corporations, and government entities.
Apply Now


