Set research direction
Translate open-ended safety questions into scoped hypotheses, decision criteria, and sequenced workstreams.
AI Safety research portfolio · Senior technical leadership
Postdoctoral Associate in AI Safety at MBZUAI
AI Safety Researcher
AI Safety researcher working on adversarial evaluation, multimodal alignment, model robustness, and scalable safety methodologies for generative AI systems.
Executive profile
A senior research profile built around rigorous evaluation and clear safety decisions.
Samuele is a Postdoctoral Associate at MBZUAI with a PhD in Artificial Intelligence focused on Responsible AI for vision and language. His research spans AI safety, security, robustness, constitutional AI, multilingual safety, multimodal systems, and model evaluation.
His work turns safety questions into testable methodologies: benchmarks, evaluation frameworks, adversarial tests, datasets, and experimental pipelines. He supervises MSc and PhD researchers, coordinates collaborative projects, reviews experimental quality, and translates technical findings into decision-relevant recommendations.
Core expertise
Technical evaluation, system security, and research leadership—without subjective proficiency scores.
Selected research
Six case studies spanning post-alignment model drift, multimodal safeguards, multilingual robustness, and AI security. Open a study for the complete research frame.
Harmless fine-tuning can partially undo alignment, revive unlearned capabilities, or restore latent behaviours acquired earlier in training.
Benign post-alignment updates may reintroduce harmful behaviour while preserving downstream task performance, making safety erosion difficult to detect.
First author on the work connecting fine-tuning reversion to training-history geometry and evaluating v_rev as a causal mediator of post-alignment drift.
Representational drift rapidly aligns with v_rev; selectively blocking that direction changes final alignment and reduces harmfulness with little task cost in the evaluated setup.
Provides a mechanistic signal and intervention target for maintaining safety when aligned models are customized after deployment.
Co-authored the study in a focused two-researcher collaboration linking mechanistic analysis to deployment safety decisions.
Safety controls in text-to-image systems can weaken after ordinary fine-tuning, even when the downstream task and data are benign.
A model may pass its original safety evaluation but become less reliable after customization, while still appearing useful on standard quality measures.
Co-authored the work and advised project development on benchmark design, experimentation, and publication for harmful-content mitigation in text-to-image models.
The benchmark shows why safety alignment must be re-evaluated after benign adaptation rather than treated as a fixed pre-deployment property.
Gives safety teams a deployment-aware framework for deciding whether a customized generative model remains safe enough to release.
Provided research mentorship on design, implementation, experimental review, and publication strategy.
Vision-language representations can preserve unsafe concept associations that propagate into downstream retrieval and generation systems.
Representation-level associations can make harmful visual-textual content difficult to filter consistently without damaging useful model behaviour.
First author; led the research on safety-aware representation learning, experimental evaluation, and technical communication.
Safe-CLIP reduces unsafe concept associations in CLIP while preserving useful vision-language behaviour across evaluated tasks.
Connects representation-level intervention to practical guardrails, content-safety evaluation, and multimodal model assurance.
Coordinated the research with collaborators at AImageLab and contributed to the open-source code, dataset, and project release.
Safety behaviour that is reliable in one language may not transfer consistently across languages or survive later fine-tuning.
Uneven safeguards create cross-lingual attack surfaces in globally deployed generative AI systems.
Conducted research on multilingual LLM safety, red teaming, and Llama models as a Research Scientist Intern within Meta GenAI Trust & Safety.
The publicly disclosed research shows that fine-tuning attacks in one language can degrade multilingual safety alignment.
Addresses a core assurance challenge for AI systems expected to operate safely across global user populations.
Collaborated with an industrial research team while translating an open safety question into reproducible cross-lingual evaluation.
Fine-tuning can alter established safety behaviour after a model leaves its original training environment.
Adaptation may create multilingual safety degradation that standard pre-deployment tests do not capture.
First author; led the study of multilingual fine-tuning attacks, experimental analysis, and communication of the resulting safety implications.
Fine-tuning attacks in one language can break safety alignment across languages, suggesting that parts of the relevant safety information are language-agnostic.
Shows why safety assurance must continue after model adaptation and account explicitly for multilingual risk.
Coordinated a cross-institution collaboration spanning Meta GenAI and academic research partners.
Security mechanisms can fail under model modification, adaptive attacks, or threat models that differ from the original evaluation.
Watermarks, monitors, and authenticity signals may be bypassed or forged if they are evaluated only under average-case conditions.
Co-authored work on activation watermarking and randomized-key defenses, and mentored researchers working on watermarking and deepfake analysis.
The portfolio treats robustness under adaptation and adaptive attack as part of the security claim, not as a secondary check.
Brings an explicit adversary and evidence model to safety monitoring, provenance, and authenticity decisions.
Supported research planning, project execution, experimental review, and publication strategy across several collaborative workstreams.
Red-teaming methodology
A general research approach for evaluating generative AI systems across languages, modalities, and model versions.
Map capabilities, interfaces, users, constraints, and deployment context.
Name assets, adversaries, access, incentives, and security boundaries.
Prioritise likely misuse, safety failures, and high-consequence edge cases.
Build prompts and scenarios that probe direct, indirect, and adaptive attacks.
Curate reproducible datasets with coverage across risks and user contexts.
Combine quantitative metrics with structured qualitative review.
Compare languages, modalities, model versions, and system configurations.
Move from failure counts to vulnerability patterns and causal hypotheses.
Match controls to the failure mode, deployment layer, and residual risk.
Verify mitigations and check for regressions or displaced risk.
Translate evidence into clear findings for technical and senior stakeholders.
General research methodology for evaluating generative AI systems.
Leadership & delivery
The emphasis is research leadership: methodology, mentoring, review, coordination, and communication. No unsupported line-management claims.
Translate open-ended safety questions into scoped hypotheses, decision criteria, and sequenced workstreams.
Define threat models, datasets, baselines, metrics, ablations, and quality gates before experiments scale.
Supervise MSc, PhD, and independent researchers across constitutional alignment, world-model jailbreaking, content safety, watermarking, and deepfake analysis.
Review experimental design, implementation, failure cases, technical writing, and publication strategy across collaborative projects.
Keep parallel research tracks aligned through clear responsibilities, review points, and shared methodological standards.
Deliver graduate lectures on Responsible AI and turn technical evidence into concise risk, limitation, and mitigation statements.
Technical evidence
Risk interpretation
Mitigation options
Stakeholder decision
Experience
A research path connecting Responsible AI foundations with current generative and multimodal safety work.
MBZUAI
Research on AI safety and security for language, vision-language, and generative models, with emphasis on alignment after adaptation and robust evidence.
Meta GenAI Trust & Safety
Industry research on multilingual LLM safety, cross-lingual transfer of safety behaviour, red teaming, and Llama models.
University of Pisa & University of Modena-Reggio Emilia
Thesis: Responsible AI in Vision and Language: Ensuring Safety, Ethics, and Transparency in Modern Models. Completed with honors.
Publications
Selected peer-reviewed papers and current preprints across AI safety, security, and multimodal systems.
A deployment-aware benchmark for testing whether text-to-image safety alignment survives benign fine-tuning and downstream customization.
A geometric account of why post-alignment fine-tuning can pull model behaviour back toward earlier training-history manifolds.
Studies failure modes in multi-turn reasoning and the gap between internal reasoning signals and final model behaviour.
Investigates activation watermarking as an internal safety-monitoring signal designed to remain informative under model change.
Shows that attacks introduced through fine-tuning in one language can degrade safety alignment across multiple languages.
Removes unsafe concept associations from CLIP while preserving useful vision-language behaviour.
Talks, service & engagement
Invited lectures, conference activity, peer review, and community-building for trustworthy AI.
Invited seminar
Symposium on Security in the Age of AI, MBZUAI · February 2026
Graduate lecture
National PhD School in AI, Scuola Normale Superiore · May 2025
Conference publication
NAACL Findings · 2025
Professional service
ACM MM, CVPR, ECCV, ICCV, BMVC · 2023 - 2025
Reviewer for major international venues
Organizing committee
ECCV 2024 and MBZUAI 2026 · 2024 - 2026
Workshop organization on trustworthy and ethical machine learning
Contact
Available for conversations about AI safety research, red teaming, multimodal evaluation, and research leadership.