Prompt-injection severity in agentic systems — a working framework
draftNotes from recent red-team work: a severity model that accounts for tool use, blast radius, and authorization context, not just whether the model can be made to swear.
I'm a penetration tester and AI red teamer who specializes in cloud security. I work across AWS infrastructure, APIs, and LLM systems, finding the failure modes that survive code review and security scans, and writing the reports engineers actually use to fix them. I lead development at Beruni, an AWS posture scanner built for startups.
Two years of hands-on engagement work, with stops at Positive Technologies in Moscow, Hackers Academy, and Thincscorp. B.S. Cybersecurity from FAST NUCES Islamabad, 2026. I care about repeatable methodology, clear reporting, and findings that survive the handoff to engineering.
Lead development of Beruni, a hosted AWS security-posture scanner built for startups. Design the cloud architecture and scanning pipeline, turning AWS configuration checks and compliance mapping (SOC 2 / PCI evidence) into a repeatable workflow. Beruni replaces the manual, screenshot-by-screenshot verification auditors used to spend weeks doing by hand with automated scans and a ready-made evidence trail.
Worked on metaverse technologies and developed an interactive virtual world on Mitoworld.io using diverse digital assets, including images, videos, and 3D models. Designed and developed a presentation of the festivals and cultural heritage of KPK, Pakistan, bringing the region's stories into an immersive virtual space. The team secured 2nd Place in the competition.
Adversarial testing of LLMs and autonomous agents: prompt injection, chain-of-thought leakage, tool-use abuse, and authorization bypass. Built deliberately vulnerable agents and Python fuzzing harnesses; documented each finding with reproducible PoCs, threat models, and developer-facing remediation.
Full-scope web and internal network engagements. External and internal reconnaissance, privilege escalation, Active Directory exploitation, and post-exploitation. Tooling across Burp Suite, Metasploit, Nmap, Hydra, and Nessus. Delivered client-facing reports with CVSS-scored findings and prioritized remediation paths.
OSINT and threat-intelligence work feeding incident response; investigation cycle time reduced ~30% through better tooling and triage playbooks. Maintained the organizational risk register and supported malware and digital-forensics investigations.
Coursework across cryptography, network security, secure SDLC, and applied AI. Capstone work fed directly into AutoCSPM and the prioritization prototype.
A collaborative virtual world built during the Project-Based Learning Summer Camp at UTeM, blending Pakistani cultural heritage with spatial design. The experience featured a custom KPK avatar, a Bab-e-Khyber gateway, and a traditional Hujra lounge with Persian rugs, prayer mats, and floor seating. When spatial barrier physics could not block the upper level, a playful "Police Cat" stop-sign wall became the improvised access control. The project secured 2nd Place in the competition.
A dual-mode cloud security posture scanner: pre-deployment against infrastructure-as-code, and live against deployed AWS accounts. Maps findings to CIS Benchmarks and OWASP cloud guidance, and outputs evidence formatted for compliance review rather than a thousand-line CSV. MVP shipped; multi-cloud and policy-as-code on the roadmap.
A Spring Boot training lab that implements the OWASP API Top 10 as runnable vulnerabilities: BOLA, broken authentication, mass assignment, injection. Parallel vulnerable and fixed branches let teams diff each exploit against its patch, line by line.
A prototype that ranks vulnerabilities by exploitability and environmental fit, pulling CVSS, EPSS, and CISA KEV signals with ML-generated remediation suggestions. Built for the “ten thousand highs” problem, where every finding is a P1 and triage collapses under its own volume.
Written for the Beruni blog: the ten AWS misconfigurations that show up most across IAM, S3, EC2, VPC, and RDS in real beta-account scans, each paired with a CLI remediation that doesn't need downtime.
A technical analysis of how leetspeak encoding can be used to extract system prompts from LLMs, discussing the underlying mechanisms and significant business implications for AI security.
An indepth guide on how to imply advanced techniques that involves Authority Framing, Continuation Bootstrapping and much more techniques that helps in Understanding the core funcionality of LLMs and how they can be taken Advantage of to bypass the security of the system.
Notes from recent red-team work: a severity model that accounts for tool use, blast radius, and authorization context, not just whether the model can be made to swear.
The three shapes of broken object-level authorization I see in nearly every engagement, what they share, and the test cases that catch them without false positives.
A walkthrough of the scanner architecture, the CIS rule engine, and why “compliance evidence” is the only output that ever changes behaviour in mid-sized orgs.
More posts as engagements close and material becomes shareable. Most of my output lives in client reports rather than public writing.
Email is the fastest channel; I reply within a working day. Currently open to remote pentest and AI red team roles, and selective freelance engagements.
CV available on contact.
By default, I don't know. Someday it's one of these, someday it's all of them at once, someday it's none. In short: trying to navigate life alongside its metaphysical realities.
Press M then I then K — anywhere on this page. No mouse needed.
A full-screen navigation shortcut hidden inside the site. Old-school keyboard ritual. Type the initials.