Education & Research

Robotics and computer science, now AI safety and alignment

A doctorate in AI safety and alignment, a masters in computer science and a bachelors in robotics, alongside two working papers. The thread running through the research is evaluation validity, whether a benchmark measures the capability it claims to, and what follows when it does not.

Doctorate
AI safety and alignment
Working papers
2, in preparation
Degrees
Ph.D., M.Sc., B.Tech

Education

Robotics, then computer science, then AI safety

The route into this work runs from control systems and embedded engineering, through software and data platforms, to alignment research. Each step is visible in what I build now.

  1. Ph.D., Artificial Intelligence & Machine Learning

    In progress

    Paris American International University, France

    AI safety and alignment.

    Attenuation of Safety Alignment Across Language and Modality in Multimodal LLMs

  2. M.Sc., Computer Science

    First Class Distinction

    University of East London, United Kingdom

    AI, computer vision, advanced software engineering and big data analytics.

    The software engineering half of the work, from system design through to the data platforms underneath.

  3. B.Tech, Robotics & Automation Engineering

    First Class Distinction

    Parul University, India

    Robotics, control systems and embedded engineering.

    Where the perception and on-device inference work comes from, including the drone vision and firefighting robot projects.

Doctoral research

Attenuation of Safety Alignment Across Language and Modality in Multimodal LLMs

Ph.D., Artificial Intelligence & Machine Learning

Multimodal Foundation ModelsRLHF & Preference OptimizationAI Safety & Alignment

The question

How much of the apparent cross-modal safety gap is a reading-ability artifact rather than an alignment gap?

What the field reports

Multimodal safety benchmarks report that guardrails weaken when a harmful request arrives as an image rather than as text, or in a lower-resource language rather than in English. That result is usually read as an alignment failure: the safety training did not transfer.

Why that reading may be wrong

The measurement cannot distinguish two very different models. One understood the request and refused. The other could not parse it at all. Both score as a refusal, and both score as a safety pass. A model that simply cannot read Swahili looks as aligned as one that read the request and declined it. The apparent gap may be measuring comprehension, not alignment.

The contribution

A comprehension control run alongside the safety evaluation. Before a refusal counts as evidence of alignment, the model has to demonstrate on a benign task that it could read the input at all. Separating the two lets the cross-modal gap be split into the part that is an alignment failure and the part that is a perception artifact.

Method

  • Parallel prompt sets across English, Swahili and Arabic, with translation quality stated rather than assumed
  • The same prompts as image-embedded text, so language and modality vary independently
  • A benign comprehension control per prompt, per language and per modality
  • A benign over-refusal control, so a model that refuses everything does not score as safe
  • Attack success rates with bootstrap intervals, and judge agreement reported as kappa against human labels on a verified subset

Status

Active research. The instrument and the controls are designed; the measurement has not been run, so no results are reported.

Release policy

Prompts come only from already-published benchmarks, with no novel attacks. No working jailbreak strings are committed, and results are reported as aggregate rates rather than per-prompt transcripts.

Working papers

Two preprints in preparation

Both are in preparation rather than published, and are labelled as such here for the same reason every figure on this site cites its source.

  • 01Preprint in preparationarXiv, 2026

    Multi-Agent Orchestration with Model Context Protocol: A Framework for Reliable Tool-Augmented LLM Systems

    Failure modes of tool-augmented agents measured per topology: loops, tool misselection, error cascades and context exhaustion, each with a reproduction seed and a rate. The finding that a supervisor topology can cost several times the tokens for no accuracy gain came out of this work.

    mcp-server-and-agent
  • 02Preprint in preparationarXiv, 2026

    Safety Alignment Across Language and Modality: Perception Confounds in Multimodal LLM Guardrail Transfer

    The dissertation's central argument: that cross-modal and cross-lingual safety benchmarks conflate refusal with comprehension, and that a comprehension control is required before a refusal rate can be read as an alignment result.

    Instrument and controls designed; measurement not yet run.

Agenda

What I am working on beyond the dissertation

  • Evaluation validity

    Whether a benchmark measures the capability it claims to. Judge calibration, comprehension controls, contamination, and the smallest delta a sample size can actually detect.

  • Sycophancy & reward hacking

    Preference-trained models learn to agree rather than to be right, because agreement is what the reward signal rewards. Measuring sycophancy separately from accuracy, and finding where an objective is being gamed rather than satisfied.

  • Safety under distribution shift

    How alignment behaves away from the language and modality it was trained on, and how much of the observed change is a measurement artifact rather than a real gap.

  • Agent reliability

    Failure modes of tool-using systems as a measurable property of topology, covering loops, misselection, error cascades and context exhaustion, and where human review belongs in an automated pipeline.

  • Eval gaming & contamination

    Benchmarks lose meaning once they are in the training set. Decontamination that is checked rather than assumed, and held-out construction that survives the next scrape.

  • Human-in-the-loop thresholds

    Confidence gating as a design problem: where the threshold goes, what it costs per class, and how to tell an abstention apart from a failure.

Contact

Get in touch

Happy to talk about anything here, or about LLM systems, evaluation and serving generally. Mentioning a repository by name gets you a faster and more useful answer.

Based in
Abu Dhabi, UAE