Alizishaan Khatri
发表论文 6 篇 · 总被引 1 次 · h-index 1
代表论文
- Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families (2026 · 被引 1)
- Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models (2026 · 2026 56th Annual IEEE International Conference on Dependable Systems and Networks Workshops (DSN-W) · 被引 1)
- What's New in TensorFlow 2.0 (2019 · 被引 1)
- Preventing overfitting in deep learning using differential privacy (2026 · arXiv.org)
- Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations (2026)
- Latent Space Probing for Adult Content Detection in Video Generative Models (2026 · 2026 56th Annual IEEE International Conference on Dependable Systems and Networks Workshops (DSN-W))