> For the complete documentation index, see [llms.txt](https://yanyun-wangs-gitbook.gitbook.io/yanyun-wangs-gitbook/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://yanyun-wangs-gitbook.gitbook.io/yanyun-wangs-gitbook/reading-notes/reading-note-threats-to-pre-trained-language-models-survey-and-taxonomy.md).

# Reading Note: "Threats to Pre-trained Language Models: Survey and Taxonomy"

Guo, Shangwei, et al. "Threats to pre-trained language models: Survey and taxonomy." arXiv preprint arXiv:2202.06862 (2022).

## Abstract

**Pre-trained language models (PTLMs)** have achieved great success and remarkable performance, while there are growing concerns regarding their security issues.

**Reasons** that make PTLMs particularly vulnerable:

* Threats can occur at <mark style="background-color:purple;">different stages</mark> of PTLM pipeline (pre-training, finetuning, inferring) raised by <mark style="background-color:purple;">different malicious</mark> entities (model publisher, downstream service provider, user);
* Two types of <mark style="color:orange;">model transferability</mark> facilitate attacks (landscape & portrait);
* Four categories of attacks based on different attack goals (integrity threats: backdoor and evasion attacks & privacy violations: data and model).

<figure><img src="https://725511345-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fo36vbcqVTETuSOVJefc3%2Fuploads%2FVPRTaiuAbGVFBW9HvG1L%2Fimage.png?alt=media&amp;token=d6bd8b95-6d81-41f0-aa4e-a721e0d017c1" alt=""><figcaption><p><strong>The PTLM system pipeline with possible attack goals enabled by two types of transferability.</strong></p></figcaption></figure>

## 1 System and Threat Overview

### 1.1 Two types of existing PTLMs

* <mark style="color:purple;">Autoencoding Model (AE)</mark>: pre-trained through corrupting input tokens and attempting to reconstruct the original sentences (e.g., next sentence prediction, masked language model -> <mark style="background-color:purple;">BERT</mark>: pre-train deep bidirectional representations, some input tokens are replaced by \[MASK]).
* <mark style="color:purple;">Autoregressive Model (AR)</mark>: trained to encode unidirectional context and predict the token of current time-step according to the tokens read before (e.g., text generation -> <mark style="background-color:purple;">GPT</mark>).

### 1.2 PTLM Pipeline

* *Pre-training*: Model Publisher (MP) trains a foundation PTLM from enormous <mark style="background-color:purple;">unsupervised</mark> corpus.
* *Fine-tuning*: Downstream Service Provider (DSP) obtains PTLM from MP, and transfers it to a specific downstream model (usually append an auxiliary structure such as a linear classifier to PTLM, and fine-tune with downstream corpus in a <mark style="background-color:purple;">supervised</mark> manner).
* *Inferring*: DSP deploys the fine-tuned model as a NLP service, and provides APIs for users. When receiving text queries, the inference system conducts forward propagation to obtain the output.

### 1.3 Attack Goals

**Integrity Attacks**: to compromise the integrity of model <mark style="background-color:purple;">parameters</mark> or <mark style="background-color:purple;">predictions</mark>.

* *Backdoor*: by malicious <mark style="background-color:orange;">MP</mark>, embed backdoors into PTLM, which can be activated by malicious input (containing specific <mark style="color:orange;">triggers</mark>) of the downstream model.
* *Evasion attack*: by malicious <mark style="background-color:orange;">user</mark> at inferring time, craft <mark style="color:red;">adversarial examples</mark> to mislead the downstream model to produce wrong results.

**Privacy Attacks**: to steal <mark style="background-color:purple;">sensitive information</mark> from pre-trained or downstream models.

* *For data*: by <mark style="background-color:orange;">DSP or user</mark>, recover attributes, keywords, or entire sentence of training corpus.
* *For model*: by <mark style="background-color:orange;">user</mark>, extract the proprietary pre-trained model.

### 1.4 Model Transferability

* *Landscape Transferability*: downstream models from the same PTLM share similar language representation features, transfer attack <mark style="color:orange;">between</mark> them.
* *Portrait Transferability*: inject backdoors into PTLM, <mark style="color:orange;">inherited</mark> by downstream models.

<figure><img src="https://725511345-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fo36vbcqVTETuSOVJefc3%2Fuploads%2FYwvr1RPGR4YdQGKJHxMt%2Fimage.png?alt=media&amp;token=01cb8b1d-f3ab-437e-afa2-2a88ba06d9a4" alt=""><figcaption><p><strong>A list of existing attacks on PTLM systems.</strong> Position at original paper: Table 1, page 3.</p></figcaption></figure>

## 2 Integrity Threats

### 2.1 Backdoor Attacks

To inject the backdoor to the victim model by <mark style="color:orange;">poisoning the training samples</mark> or directly <mark style="color:orange;">manipulating the parameters</mark>. The infected model still behaves <mark style="background-color:purple;">normally for clean samples</mark>, but outputs wrong predictions for input containing attacker-specified <mark style="color:orange;">triggers</mark>.

Two categories based on adversary’s knowledge:

* *Task-specific attacks***:** adversarial MP has knowledge of downstream tasks (e.g., fine-tuning methods, partial or full finetuning corpus), and builds backdoored PTLMs <mark style="background-color:purple;">specifically</mark> for those tasks (e.g., RIPPLe \[1], context-aware generative model-based \[2]). <mark style="background-color:purple;">Not realistic</mark> in most cases.
* *Task-agnostic attacks*: enable the embedded backdoor to transfer to arbitrary downstream models (e.g., BadPre \[3], NeuBA \[4], POR-based \[5], layer weight poisoning training \[6]).

### 2.2 Evasion Attacks

A malicious <mark style="background-color:purple;">user</mark> designs <mark style="color:orange;">adversarial text inputs</mark>, which are semantically indistinguishable from normal ones, to mislead the target downstream models in the <mark style="background-color:purple;">inferring</mark> phase.

**2.2.1 White-box Attack**

To compute the malicious input based on the model parameters (e.g., measure the gradient distance between normal and adversarial words \[7]).

**2.2.2 Black-box Aattack**

A possible strategy is to construct a <mark style="color:orange;">shadow model</mark> from which the adversarial examples are generated (has high chance when shadow and victim models are transferred from <mark style="background-color:purple;">the same PTLM</mark>).

Two categories based on the granularity of adversarial perturbations:

***1)** Word-level attacks*

* <mark style="color:orange;">Heuristic generation</mark>: design perturbation through pre-defined rules (e.g., <mark style="color:green;">TextFooler</mark> based on word importance \[8], swarm optimization-based method \[9], <mark style="color:green;">Adv-OLM</mark> to select words for replacement \[10], population-based optimization \[11], <mark style="background-color:purple;">syntactically</mark> incorrect word generation \[12], transformer-based extension of TextFooler for high transferability \[13], population-based genetic algorithm for high transferability \[14]).
* <mark style="color:orange;">Automatic generation</mark>: leverages an additional model to automatically generate substitution words to achieve better semantic indistinguishability (e.g., <mark style="color:green;">BERT-Attack</mark> to find important words by the \[MASK] \[15], <mark style="color:green;">BAE</mark> to utilize contextual perturbations from a BERT masked language model \[16], <mark style="color:green;">CLARE</mark> through a mask-then-infill procedure \[17], modification with shared words \[18], <mark style="color:green;">MORPHEUS</mark> to perturb the inflectional morphology of words \[19]).

***2)** Sentence-level attacks*

* Craft adversarial sentences with exploitation of sentence structures and contexts, instead of replacing certain words (e.g., irrelevant sentences for machine reading comprehension \[20], <mark style="color:green;">T3</mark> with tree-based autoencoder <mark style="background-color:yellow;">embedding discrete text into a</mark> <mark style="color:red;background-color:yellow;">continuous representation space</mark> \[21], paraphrase datasets \[22, 23]).

## 3 Privacy Threats

### 3.1 Data Privacy Attacks

ML models can memorize data \[24], which allows malicious <mark style="background-color:purple;">DSP or users</mark> to steal key information of training or inference samples from <mark style="background-color:purple;">embedding codes and PTLMs</mark>.

Three categories according to the type of extracted information:

* *Embedding inversion attacks*: DSP can <mark style="color:orange;">invert</mark> the original sentence of an <mark style="background-color:orange;">inference input</mark> based on the corresponding <mark style="color:orange;">embedding code</mark> \[25].
* *Attribute inference attacks*: <mark style="color:purple;">MIA</mark> \[25] (also see another reading note [here](/yanyun-wangs-gitbook/reading-notes/reading-note-membership-inference-attacks-on-machine-learning-a-survey.md)) & keyword inference attacks (whether certain keywords exist in an unknown inference sentence) \[25, 26].
* *Corpus inference attacks*: extract the <mark style="background-color:orange;">training corpus</mark> from PTLMs \[27] or downstream models \[28].

### 3.2 Model Privacy Attacks

A malicious <mark style="background-color:purple;">user</mark> could perform <mark style="color:purple;">model extraction attacks (MEAs)</mark> to reconstruct the proprietary model by querying the system in the inferring phase.

Two categories according to the extraction goals:

* *Accuracy extraction attacks*: extract a model with <mark style="background-color:orange;">similar accuracy</mark> on the text data as the victim PTLM (e.g., task-specific query generator \[29] and algebraic extraction attack \[30], both against BERT-based models).
* *Fidelity extraction attacks*: to steal a PTLM with <mark style="background-color:orange;">similar behaviors</mark> as the victim one (e.g., imitation attack \[31] and querying gibberish data for monolingual models \[32]).

## 4 Future Directions (<mark style="color:red;">at 2022</mark>)

* **Robustness enhancement**: an arms race - between designing more sophisticated attacks against <mark style="background-color:purple;">backdoor detection or removal methods</mark> (e.g., \[33, 34]) and combining the characteristics of PTLM systems with <mark style="background-color:purple;">conventional</mark> robustness solutions (e.g., adversarial training \[35]).
* <mark style="background-color:red;">**Trade-off**</mark>**&#x20;between utility and security**: obfuscating model parameters or inference behaviors is common for preventing information leakage (e.g., added Gaussian noise to defeat MIA \[26]), but can affect the utility of PTLMs.
* **Transferability improvement**: to reduce the attack <mark style="background-color:purple;">transferability</mark> of PTLMs while maintaining its <mark style="background-color:purple;">generalization</mark>.

<mark style="background-color:yellow;">(For backdoor threats, it is valuable to devise a finetuning method that only transfers the knowledge of PTLM for normal data while forgetting the knowledge of malicious data with triggers.)</mark>

## Selected References

\[1] Keita Kurita, Paul Michel, and Graham Neubig. Weight poisoning attacks on pretrained models. In ACL, 2020.

\[2] Xinyang Zhang, Zheng Zhang, Shouling Ji, and Ting Wang. Trojaning language models for fun and profit. In S\&P, 2021.

\[3] Kangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo, Tianwei Zhang, Jiwei Li, and Chun Fan. Badpre: Task-agnostic backdoor attacks to pretrained NLP foundation models. arXiv preprint, 2021.

\[4] Zhengyan Zhang, Guangxuan Xiao, Yongwei Li, Tian Lv, Fanchao Qi, Zhiyuan Liu, Yasheng Wang, Xin Jiang, and Maosong Sun. Red alarm for pretrained models: Universal vulnerability to neuron-level backdoor attacks. In ICML, 2021.

\[5] Lujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li, Jing Chen, Jie Shi, Chengfang Fang, Jianwei Yin, and Ting Wang. Backdoor pre-trained models can transfer to all. In CCS, 2021.

\[6] Linyang Li, Demin Song, Xiaonan Li, Jiehang Zeng, Ruotian Ma, and Xipeng Qiu. Backdoor attacks on pre-trained models by layerwise weight poisoning. In EMNLP, 2021.

\[7] Yong Cheng, Lu Jiang, and Wolfgang Macherey. Robust neural machine translation with doubly adversarial inputs. In ACL, 2019.

\[8] Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. Is BERT really robust? a strong baseline for natural language attack on text classification and entailment. In AAAI, 2020.

\[9] Yuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu, Meng Zhang, Qun Liu, and Maosong Sun. Word-level textual adversarial attacking as combinatorial optimization. In ACL, 2020.

\[10] Vijit Malik, Ashwani Bhat, and Ashutosh Modi. Adv-OLM: Generating textual adversaries via OLM. arXiv preprint, 2021.

\[11] Rishabh Maheshwary, Saket Maheshwary, and Vikram Pudi. Generating natural language attacks in a hard label black box setting. In AAAI, 2021.

\[12] Fan Yin, Quanyu Long, Tao Meng, and Kai-Wei Chang. On the robustness of language encoders against grammatical errors. In ACL, 2020.

\[13] Chris Emmery, ́ Akos K ́ ad ́ ar, and Grzegorz Chrupała. Adversarial stylometry in the wild: Transferable lexical substitution attacks on author profiling. arXiv preprint, 2021.

\[14] Liping Yuan, Xiaoqing Zheng, Yi Zhou, Cho-Jui Hsieh, and Kai-Wei Chang. On the transferability of adversarial attacks against neural text classifier. In EMNLP, 2021.

\[15] Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. BERT-ATTACK: Adversarial attack against BERT using BERT. arXiv preprint, 2020.

\[16] Siddhant Garg and Goutham Ramakrishnan. BAE: BERT-based adversarial examples for text classification. arXiv preprint, 2020.

\[17] Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, and Bill Dolan. Contextualized perturbation for textual adversarial attack. arXiv preprint, 2020.

\[18] Zhouxing Shi and Minlie Huang. Robustness to modification with shared words in paraphrase identification. arXiv preprint, 2019.

\[19] Samson Tan, Shafiq Joty, Min-Yen Kan, and Richard Socher. It’s Morphin’Time! combating linguistic discrimination with inflectional perturbations. arXiv preprint, 2020.

\[20] Jieyu Lin, Jiajie Zou, and Nai Ding. Using adversarial attacks to reveal the statistical bias in machine reading comprehension models. arXiv preprint, 2021.

\[21] Boxin Wang, Hengzhi Pei, Boyuan Pan, Qian Chen, Shuohang Wang, and Bo Li. T3: Treeautoencoder constrained adversarial text generation for targeted attack. In EMNLP, 2020.

\[22] Wee Chung Gan and Hwee Tou Ng. Improving the robustness of question answering systems to question paraphrasing. In ACL, 2019.

\[23] Yuan Zhang, Jason Baldridge, and Luheng He. Paws: Paraphrase adversaries from word scrambling. In NAACL-HLT, 2019.

\[24] Vitaly Feldman and Chiyuan Zhang. What neural networks memorize and why: Discovering the long tail via influence estimation. NeurIPS, 2020

\[25] Congzheng Song and Ananth Raghunathan. Information leakage in embedding models. In CCS, 2020.

\[26] Xudong Pan, Mi Zhang, Shouling Ji, and Min Yang. Privacy risks of general-purpose language models. In S\&P, 2020.

\[27] Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. Extracting training data from large language models. In USENIX Security, 2021.

\[28] Santiago Zanella-Beguelin, Lukas Wutschitz, Shruti Tople, Victor R ̈ uhle, Andrew Paverd, Olga Ohrimenko, Boris Kopf, and Marc Brockschmidt. Analyzing information leakage of updates to natural language models. In CCS, 2020.

\[29] Xuanli He, Lingjuan Lyu, Qiongkai Xu, and Lichao Sun. Model extraction and adversarial transferability, your BERT is vulnerable! arXiv preprint, 2021.

\[30] Santiago Zanella-Beguelin, Shruti Tople, Andrew Paverd, and Boris Kopf. Grey-box extraction of natural language models. In ICML. PMLR, 2021.

\[31] Qiongkai Xu, Xuanli He, Lingjuan Lyu, Lizhen Qu, and Gholamreza Haffari. Beyond model extraction: Imitation attack for black-box NLP APIs. arXiv preprint, 2021.

\[32] Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher. The thieves on sesame street are polyglots-extracting multilingual models from monolingual APIs. In EMNLP, 2020.

\[33] Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun. ONION: A simple and effective defense against textual backdoor attacks. In EMNLP, 2021.

\[34] Chun Fan, Xiaoya Li, Yuxian Meng, Xiaofei Sun, Xiang Ao, Fei Wu, Jiwei Li, and Tianwei Zhang. Defending against backdoor attacks in natural language generation. arXiv preprint, 2021.

\[35] Xiaodong Liu, Hao Cheng, Pengcheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao. Adversarial training for large neural language models. arXiv preprint arXiv:2004.08994, 2020.
