1
Fast Multi-dimensional Refusal Subspaces via RFM-AGOP
用RFM-AGOP快速定位多维度拒绝子空间,为AI安全可解释性提供高效新方法。
arXiv:2607.02396v1 Announce Type: new Abstract: Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both saf…