Open to robotics and AI engineering roles, and PhD positions

Teaching robots to say what they mean, and to know when they can't.

I'm a robotics and AI engineer. I build ROS2 planners, embedded firmware and verification layers that stop unsafe plans before hardware moves, and the benchmarks that measure where language-model planners fail.

MSc AI & Robotics (Commendation), University of Hertfordshire. C++20, Python, ROS2, PyTorch.

start goal A goal B keep-out zone
PathLengthSure byZone
Shortest safe1.00x70%clear
Legible1.20x42%enters
Legible and safe1.06x61%clear
or drag anything on the map

Sure by: how far along the path a watcher is 90% sure of the goal. Earlier is clearer. A toy model of the trade-off legible-motion-bench measures, computed live in your browser. An illustration, not benchmark data.
2
systems built under contract: a portal that runs an institute every day, and firmware for a commercial prototype.
0
missed dangers in the LLM safety verifier across 2,071 constructed and 80 real-model cases.
589
tests and label proofs re-run on every commit of plan-failure-bench.
32/32
certified legibility bounds hold with zero violations. Submitted to IEEE RA-L.

Engineering

A portal that an institute runs on every day, firmware for a commercial prototype, and a planning and verification stack for robots that take instructions from language models. Every project has its code, its tests and its limits on its own page.

Robotics and software

C++20, Python, ROS2 (Nav2, plugin development), Git and CI/CD, Docker

Simulation

Isaac Sim, MuJoCo, Gazebo, Unity, sim-to-real transfer

AI and perception

PyTorch, TensorFlow, TensorRT, computer vision, LLM integration

Hardware and embedded

JD humanoid, Unitree Go1, Jetson Nano, ESP32 and LoRa, bare-metal C/C++

Contract work
Safina Portal In production · React, TypeScript, Supabase

A school management system that runs Safina International day to day: seven roles, prorated billing, payroll, and an append-only audit trail. I am the sole engineer. Showcase repo

33audited tables nothing bypasses through the API
ESP32-CAM motion detector Deployed · C++, ESP32, LoRa mesh

Deterministic firmware for a commercial intrusion-detection prototype, built under contract for Muxtronics. No ML and no vision libraries anywhere. Code

16×12decision grid distilled from 800×600 frames
Robot planning and verification
ROS2 deterministic planning C++20 · Nav2 plugins · CI-verified

A* and D* Lite as Nav2 plugins over a ROS-free core, validated against Dijkstra with exact integer arithmetic across 185,237 fuzzed replans. Code

4.3xfaster mean replanning with D* Lite when an obstacle blocks the path
ROS2 LLM safety verifier ROS2 · Nav2 · CI-replayed

A deterministic gate between LLM planners and Nav2. Collision, continuity and map-bounds checks reject hallucinated trajectories before a controller sees them. Code

35/35 · 32/32unsafe plans caught from two models, zero false positives
llm-nav-shield Neurosymbolic · safe autonomy

An LLM proposes a route, a deterministic verifier checks it, and a provably correct planner recovers a safe one. When no safe path exists it halts, in 10/10 such cases, and invents nothing. Code

38/38flawed proposals recovered, none unsafe forwarded
exact-predicates C++ · numerical robustness

Geometric predicates that cannot be wrong, grown from a real D* Lite bug where two equal keys landed one ulp apart. Exactness costs about 2x here, not orders of magnitude. Code

657adversarial cases where float is wrong and exact is right
toolcall-contract LLM agents · Python

LLM tool calls that parse, type-check and validate, and still break the contract. A schema check, which is what most agent frameworks do, passes 39 of 40 of these calls. Code

32/40pass the contract

Every project is also explained without the jargon, with the same numbers and the same stated limits. A side interest in numerical analysis lives there too: a note on finite-time blowup in one-dimensional fluid models.

Research

Two lines of work: how language-model planners fail, measured against machine-checked ground truth, and how robots should move and speak so people can calibrate their trust. Every figure is backed by committed data a stranger can re-run. Academic CV (PDF)

Certified bounds on achievable legibility under a path cost budget

A benchmark measures what a planner achieved. This certifies what nothing could achieve: an upper bound on legibility over every admissible trajectory, not only the ones somebody tried. All 32 world and budget pairs hold with zero violations, and the narrowest gap anywhere is 0.0064.

Where it stops. The observer model is exactly reproducible and has never been validated against people, so these are bounds on a stated objective, not on what a human would infer.

Certified intervals, by world
door_pair
0.0269
narrow_gap
0.0456
fan_outer
0.0282
wall_choice
0.0128
fan_middle
0.0567
legibility
0.40.60.81.0
gap

Each bar runs from a trajectory that exists to a ceiling nothing can exceed. Five of the eight worlds, each at the budget where its interval is widest.

plan-failure-bench: a machine-checkable benchmark of how language model planners fail

Every instruction either admits a valid plan or plants exactly one trap, and every ground-truth label carries a proof re-verified in CI. A deterministic checker scores each response, with no human or LLM judge. Across 18 runs on four models the failures are distinct and stable: a frontier reasoning model survives full semantic obfuscation, while smaller models split between silent compliance and surface-anchored refusal.

Where it stops. Two symbolic PDDL environments and four models. Detection is never reported without its paired false positive count.

60

instructions, each with a machine-checked proof of ground truth

18

complete model runs, every record committed

589

tests and label proofs, re-run on every commit

0

human or LLM judgements in the scoring loop

Enhancing human-robot companionship: verbal and non-verbal communication cues in humanoid robots

A full sim-to-real pipeline on the JD humanoid, from an Isaac Sim digital twin through ROS2 hardware integration, built so a gestural and a spoken condition could be delivered identically to 20 participants. Speech conveyed intent significantly more clearly than gesture, while engagement and warmth held comparable across both. The study was run during the MSc (2023-2024); the manuscript was written up for journal submission in 2025.

Where it stops. One deliberately mundane scenario, the robot asking to be recharged, with 20 participants. Every gestural failure was a failure of joint attention, which is why the design target is hybrid.

Intent identified correctly
The JD humanoid robot used in the study
Speech
95%
Gesture
80%

Clarity effect d = 0.58, p = 0.017. Within-subjects, counterbalanced, RoSAS-validated, N = 20.

legible-motion-bench: do language models know when their motion is legible?

Measures legibility against path cost and constraint satisfaction, computed exactly, with no human rater and no model judge. Asked to plan in the same worlds, three language models called 116 of their 120 trajectories legible, including all 25 that were not physically possible.

Where it stops. Three models at one temperature is a pilot, not a finding. Next: whether the metric tracks what people actually perceive.

Claimed vs possible, of 40 decodes
Qwen 2.5 7B
40
26
Llama 3.3 70B
40
29
Gemini 3.6 Flash
36
40
called legible by the modelphysically possible

About

Portrait of Munawar Kazmi

I work at the boundary between robotics research and shipped engineering. My academic work asks how humanoid robots should communicate so that people trust them appropriately: not too little, not too much. My applied work is the planners, verification layers and embedded firmware those robots run on.

I hold both to one standard: a stranger should be able to reproduce every number, and every project should say where it stops.

Education
MSc AI & Robotics, CommendationUniversity of Hertfordshire, 2023-2024
BEng Mechatronics EngineeringNUST, 2018-2022
Experience
Visiting LecturerSep 2026 - Present
UCLTILS · Multan, Pakistan

Teach Computing Principles and Logic, a first-year module of the University of London BSc Artificial Intelligence taught at UCLTILS.

Full-Stack Software EngineerJul 2026 - Present
Safina International (Contract) · Multan, Pakistan

Sole engineer of the institute's live platform, in production across seven user roles and 33 audited tables.

Visiting Lecturer, AI & RoboticsAug 2025 - Present
Pak-Turk Maarif Schools · Multan, Pakistan

Designed and teach an ML curriculum for 30+ students with no prior background, taking teams through training and evaluating their own classifiers on custom tasks.

Embedded Systems EngineerJan 2025 - May 2025
Muxtronics (Contract) · Multan, Pakistan

Deterministic C++ firmware for real-time image processing on ESP32, and the LoRa mesh protocol that carried alerts off the device, for a commercial intrusion-detection prototype.

Robotics Integration SpecialistJan 2024 - May 2024
Robot House · University of Hertfordshire

Built the gesture recognition and motion planning pipeline for the JD humanoid in Unity and C++, integrated through ROS2, and deployed PyTorch non-verbal communication models onto its real-time control loop.

Research AssociateMay 2023 - Dec 2023
University of Hertfordshire

Fine-tuned and optimised TensorFlow models for real-time inference on the Unitree Go1 quadruped, to the latency budget a walking controller leaves for perception.

Embedded Systems InternMar 2022 - Apr 2022
NCRA, National Centre of Robotics & Automation · Islamabad

Optimised bare-metal C firmware for embedded controllers at the national robotics centre, validated across 100+ safety-critical test cycles.

References

“He actively sought feedback and incorporated it into his work, demonstrating a commitment to continuous improvement. His innovative approach, coupled with a strong foundation in research and development, positions him as a valuable asset to academic or professional institutions.”
Dr Patrick HolthausReader in Interactive Assistive Technology, University of Hertfordshire. MSc supervisor.
“Munawar demonstrated exceptional skill in embedded AI deployment during his contract at Muxtronics, leading the ESP32-CAM prototype from concept to a working product.”
Umer Waheed BukhariCEO, Muxtronics
“Money cannot be quietly deleted: a wrong payment is corrected by a recorded reversal, so the record of what happened survives. And every change anyone makes is logged automatically. He proposed both of these himself. I did not ask for them, and I did not know to.”
Irum MoinPrincipal, Safina International. Full reference (PDF)
Talks and teaching
Munawar Kazmi presenting at the University of Hertfordshire
AI & Robotics in Education Conference Talk · University of Hertfordshire

Presented on HRI and edge AI applications to students and faculty, with live robot demonstrations of the companionship research.

Munawar Kazmi leading a hands-on robotics and ML workshop at Pak-Turk Maarif Schools
Hands-on robotics and ML workshop Workshop · Pak-Turk Maarif Schools

Led practical sessions covering gesture control, neural networks and real-time AI deployment. Teams built and trained working classifiers during the session.

Contact

Hiring teams

I am open to robotics and AI engineering roles in the UK, EU and USA. I am based in Pakistan and would need visa sponsorship in each.

Supervisors and research groups

I am pursuing a PhD in human-robot interaction and AI reliability, and I'm happy to share full research proposals.