Skip to main content
AI Model Governance Specialist
Functional Area
Risk
Experience
9 - 12 Years
Openings
1
Location
Mumbai
Date of Posting
15/09/26
Closing Date
30/10/26
Key Requirement
Strong AI/ML expertise with hand-on experience in RAG, LLM Evaluation, model risk controls, and enterprise AI platforms
Apply Now
Job Summary
The role will be responsible for independently evaluating AI/LLM systems, identifying potential risk vectors, designing robust evaluation frameworks, and establishing controls to support responsible and secure enterprise AI deployment.

Job Role and Responsibilities

Roles & Responsibilities

  • Conduct RAG & GenAI risk assessments across end-to-end RAG pipelines, including knowledge-base ingestion, chunking, embeddings, vector retrieval, context grounding, and generation.
  • Identify AI/LLM risks such as data leakage, prompt injection, context contamination, retrieval failures, and hallucinations.
  • Design and execute LLM evaluation and benchmarking frameworks to assess hallucination rates, toxicity, robustness, model calibration, fairness, and bias.
  • Curate golden/reference datasets and design evaluation test cases with reference answers.
  • Implement and validate LLM-as-a-Judge evaluation pipelines.
  • Establish AI governance and risk mitigation controls covering data preprocessing, pre-deployment validation gates, and post-deployment observability.
  • Independently inspect, evaluate, and validate AI/LLM model outputs.
  • Conduct technical audits of AI/LLM systems and assess model performance against defined evaluation criteria.
  • Support the development of enterprise AI control frameworks covering data provenance, model sign-off, rate-limiting, and continuous drift monitoring.
  • Evaluate automated AI assessment and hallucination detection tooling, including Ragas, TruLens, and DeepEval.

Qualifications

  • B.Tech / M.Tech in Computer Science, Quantitative disciplines, or MCA.
  • Strong quantitative and/or Computer Science foundation.
  • Thorough understanding of AI/ML taxonomy and modern Generative AI workflows.
  • Deep understanding of RAG architecture and associated risks, including vector search, embeddings, relevance scoring, and context grounding/grounding boundaries.
  • Hands-on familiarity with reference-based testing, LLM-as-a-Judge frameworks, red-teaming fundamentals, and semantic similarity evaluation.

Experience & Skills

Required Technical Skills:

  • Strong expertise in Generative AI / GenAI and Large Language Models (LLMs).
  • Strong understanding of Retrieval-Augmented Generation (RAG) pipelines.
  • Experience in LLM evaluation, model validation, AI risk assessment, and AI governance.
  • Understanding of AI risk vectors including prompt injection, hallucination, data leakage, and retrieval failures.
  • Familiarity with enterprise LLM orchestration and deployment platforms such as Amazon Bedrock, Azure OpenAI Service, and Vertex AI.

 

Preferred / Good-to-Have Skills:

  • Working knowledge of Python for querying APIs, parsing JSON outputs, and independently auditing model evaluation scripts.
  • Practical experience with NLP evaluation metrics such as ROUGE, BLEU, BERTScore, and semantic distance metrics.
  • Ability to construct enterprise AI control frameworks covering data provenance, model sign-off, rate-limiting, and continuous drift monitoring.
  • Familiarity with automated hallucination detection and AI evaluation tools such as Ragas, TruLens, and DeepEval.
Responsible Disclosure: In case you discover any security bug or vulnerability on our platform, please report it to ciso@icicisecurities.com to help us strengthen our cyber security
© 2026 I­C­I­C­I Securities. All rights reserved.