Abu Dhabi, UAEPh.D. Researcher, AI & Machine Learning

Malcom Mudhungwaza

I build and ship production AI systems, while researching how the models underneath them behave and where they break.

Agentic AI, RAG, evals and LLM serving, designed end to end and measured against what the system is actually for. My doctoral research studies how safety alignment holds up across language and modality.

  • Senior AI Engineer
    GenAI, LLM & agentic systems
  • Ph.D. Researcher
    AI safety & alignment
  • 5+ years
    production LLM & ML systems
Malcom Mudhungwaza

Malcom Mudhungwaza

Senior AI Engineer · Abu Dhabi, UAE

In practice

What that looks like day to day

Explore the domains →

Models

Open weights and hosted, chosen per requirement

Most systems here run on open-weights models, adapted and self-hosted so the behaviour, the cost and the data path stay under my control. Hosted models are used where a managed frontier model is the better answer.

Open weights

  • Qwen
  • GLM
  • DeepSeek
  • Llama
  • Mistral
  • Falcon
  • Gemma

Fine-tuned and self-hosted. Adapted to a domain with supervised and preference training, then served with quantisation against a latency target.

Hosted models

  • Anthropic Claude
  • OpenAI GPT
  • AWS Bedrock
  • Azure OpenAI

Served through cloud APIs where a managed frontier model is the right call, with the same evaluation and cost discipline applied to both routes.

What I build

Systems, not demos

The things I am usually brought in to design and ship. Each has a version running against real traffic, and a version published here as a project.

  • Agentic document pipelines

    Multimodal extraction into structured output, LLM reasoning over the result, confidence gating per document class, and a human review path for the cases that should not be automated.

  • RAG systems

    Hybrid dense and lexical search with re-ranking, chunking chosen against the document shape, and retrieval quality measured separately from generation quality.

  • LLM serving & inference optimisation

    Quantised models behind an inference gateway, sized against a stated latency target and a cost per million tokens, with routing and per-tenant budgets.

  • Post-training & domain adaptation

    Supervised fine-tuning and preference optimisation on open-weights models, with the objective chosen against the feedback actually available rather than the newest paper.

  • Evals & CI regression gates

    Golden datasets, LLM-as-judge scoring calibrated against human labels, and regression gates that fail a build when quality drops by a detectable margin.

  • Voice agents & multimodal

    Streaming ASR into dialogue orchestration into speech synthesis, budgeted per hop against time-to-first-audio rather than total generation time.

  • Agentic AI & MCP servers

    MCP servers exposing internal systems as tools, orchestration with checkpointing that survives a restart, and topology chosen from the failure mode that matters most.

  • Platform layers & observability

    Async services with bounded concurrency, retry classification at the transport boundary, caching with a measured hit rate, and schemas read against their query plans.

Research

A doctorate on where safety alignment breaks

Ph.D., Artificial Intelligence & Machine Learning, in progress. The dissertation asks a question that current benchmarks cannot answer:

How much of the apparent cross-modal safety gap is a reading-ability artifact rather than an alignment gap?

A model that cannot read a request scores the same as one that read it and refused. Separating the two is the contribution.

Field
AI safety and alignment
Working papers
Multi-Agent Orchestration with Model Context Protocol: A Framework for Reliable Tool-Augmented LLM SystemsPreprint in preparation · arXiv, 2026Safety Alignment Across Language and Modality: Perception Confounds in Multimodal LLM Guardrail TransferPreprint in preparation · arXiv, 2026
Projects
26 repositories · 1128 tests

In the field

Where the work gets shown

Conferences, exhibitions and industry events across the UAE and East Africa, presenting the systems and meeting the people who run them.

  • Malcom Mudhungwaza presenting at a healthcare digital economy event

    Healthcare digital economy

    Ministry of ICT · Kenya

  • Exhibitor badge for Make It In The Emirates, ADNEC Abu Dhabi

    Make It In The Emirates

    ADNEC · Abu Dhabi

  • Malcom Mudhungwaza at a technology exhibition in Abu Dhabi

    Technology exhibition

    IHC · Abu Dhabi

Built with

  • PyTorch
  • Transformers
  • TRL
  • vLLM
  • AutoAWQ
  • LangGraph
  • FastAPI
  • PostgreSQL
  • pgvector
  • Redis
  • FAISS
  • sentence-transformers
  • rank_bm25
  • ONNX Runtime
  • faster-whisper
  • piper
  • OpenCV
  • Ultralytics
  • NumPy
  • SciPy
  • Matplotlib
  • Pydantic
  • httpx
  • anyio
  • arq
  • OpenTelemetry
  • pytest
  • datasketch
  • fastText
  • C++
  • Python
  • WebSockets

Contact

Get in touch

Happy to talk about anything here, or about LLM systems, evaluation and serving generally. Mentioning a repository by name gets you a faster and more useful answer.

Based in
Abu Dhabi, UAE