scieee AI-readable full text Open interactive document viewer

Teaching translation with AI: Bridging theory and practice through prompt engineering

Yamada, Masaru

Abstract

This chapter explores the innovative application of large language models (LLMs) in translator training, focusing on the use of few-shot prompts and chain-of-thought prompting. It proposes a novel approach that integrates metalanguages and concepts of Translation Studies into prompt engineering, moving beyond traditional natural language processing goals of improving machine translation quality. The chapter demonstrates how this method can create interactive and engaging learning experiences for translation students, allowing them to explore various translation strategies and develop critical thinking skills. Through concrete examples, the chapter illustrates the potential of LLMs to generate diverse translation variations and provide insightful analyses of translation processes. While acknowledging limitations and the need for critical evaluation, the research emphasises the positive and proactive possibilities of LLMs in translator training. This approach not only bridges the gap between translation theory and practice but also opens new avenues for autonomous learning and the development of essential skills for future translators in the AI era.

Full text

Chapter 5 Teaching translation with AI: Bridging theory and practice through prompt engineering Masaru Yamada Rikkyo University, Japan This chapter explores the innovative application of large language models (LLMs) in translator training, focusing on the use of few-shot prompts and chain-ofthought prompting. It proposes a novel approach that integrates metalanguages and concepts of Translation Studies into prompt engineering, moving beyond traditional natural language processing goals of improving machine translation quality. The chapter demonstrates how this method can create interactive and engaging learning experiences for translation students, allowing them to explore various translation strategies and develop critical thinking skills. Through concrete examples, the chapter illustrates the potential of LLMs to generate diverse translation variations and provide insightful analyses of translation processes. While acknowledging limitations and the need for critical evaluation, the research emphasises the positive and proactive possibilities of LLMs in translator training. This approach not only bridges the gap between translation theory and practice but also opens new avenues for autonomous learning and the development of essential skills for future translators in the AI era. 1 Introduction Recent advancements in artificial intelligence (AI), particularly the emergence of Large Language Models (LLMs), have generated extensive debate regarding their benefits and drawbacks. In response, the UK’s Institute of Translation and Interpreting (ITI) has articulated the Slow Translation Manifesto (ITI 2024). Drawing Masaru Yamada. 2026. Teaching translation with AI: Bridging theory and practice through prompt engineering. In JC Penet, Joss Moorkens & Masaru Yamada (eds.), Teaching translation in the age of generative AI: New paradigm, new learning?, 87–104. Berlin: Language Science Press. DOI: 10.5281/zenodo.17641072 Masaru Yamada on an analogy between fast food and AI-generated translations, this manifesto highlights the potential societal and professional harms of prioritising speed over quality in translation practices. It advocates for a renewed appreciation of slow, human-led translation, emphasising its artistry, ethical rigour, and cross-cultural competence – qualities that are often sacrificed in rapid, machine-driven translation processes. The Société française des traducteurs (SFT), the French translators’ association, has also raised concerns about the inadequate quality of AI-based translations and the transparency of the resources used for machine learning in LLMs. The SFT warns that unfettered access to AI tools, such as ChatGPT, could lead to a decline in respect for language professionals or result in an “unchecked rush of translations that we are asked to enhance” (Slator 2024). Similar statements regarding the potential crisis posed by AI technology in translation have been issued by other organisations, including CEATL (European Council of Literary Translators’ Associations 2024), ATA (American Translators Association 2024), and JAT (Japan Association of Translators 2024). The EU’s Language in the Human-Machine Era (LITHME) project, focusing on language in the technology era, aims to “prepare language researchers for what is coming” and “facilitate longer-term dialogue between linguists and technology developers” (LITHME Project 2021; see also Sayers et al. 2021). This comprehensive project compiles expert opinions and views, aiming to counteract the potential overreliance on AI-based technology and the market-driven overselling of AI (Sayers et al. 2021, Way 2025). By promoting such perspectives, the project plays a crucial role in encouraging a more cautious and thoughtful approach towards AI advancements. These investigations and statements from this project serve as a reminder to pause and critically reflect on the rapidly accelerating trends in our society. However, upon closer examination of these statements and surveys, it becomes apparent that the understanding and predictions about technology are not always accurate or up-to-date, and may not be entirely evidence-based. For instance, while LITHME’s 2021 report (Sayers et al. 2021) does touch on concepts relevant to LLMs and Generative AI (GenAI), these references are minimal and do not encompass the transformative developments that have since occurred, such as chat-based LLMs. This highlights the challenge of keeping pace with the rapid evolution of technology, which even at that time was difficult to fully anticipate. Similarly, the SFT states that the output of machine translation (MT) remains unreadable in its raw state and requires human correction through post-editing. However, it also adds that “70 percent of our member translators who responded to our survey considered PE (and by extension AI) a threat to their profession” 88 5 Teaching translation with AI (Slator 2024). The conflicting statement leaves unclear which era or type of MT it refers to, and the apparent emotional reaction reflecting fears of job loss suggests a lack of clarity about the intended scope and technological context of the claim. This also implies that such statements may not fully reflect the current state of cutting-edge technology. In the context of translator training, which is the primary focus of this book, there may not be as strong a backlash against AI as in parts of the industry. For instance, while some industry stakeholders see AI as a tool to enhance efficiency and provide opportunities for post-editing work, many independent professional translators express significant concerns. Fears about the indiscriminate use of AI, especially by untrained users, persist due to its potential to cause quality issues and undermine the value and compensation for human translators. According to ELIS (2025), around 64% of students in translation programmes report using GenAI at least occasionally, with 19% using it regularly. University staff estimate MT use at 63%, compared to 58% among students. These figures mark a steep rise from the limited adoption reported in 2024 and highlight the contrast between the comparatively cautious professional industry and the rapid integration of GenAI in translator training. This chapter proposes leveraging the accumulated knowledge assets of professional translators and translation researchers to actively interact with LLMs through basic prompts such as Chain of Thought (CoT) prompting and few-shot prompts. For example, translation memories, which are repositories of past human translations paired with their source texts, can be effectively used for translator training in conjunction with few-shot prompts. Additionally, the concepts of translation briefs and work instructions, traditionally used in professional translation, can be directly applied as prompts. Furthermore, the classical assets of translation research that describe translation strategies can be adapted for use as CoT prompting. The concepts studied in translation research, which explain the act of translation, can be collectively referred to as the metalanguages of translation (Miyata et al. 2023). These metalanguages can be effectively applied to prompt creation. These attempts are not merely an engineering perspective to enhance the accuracy of translation outputs from LLMs but are proposed as significant considerations for translator training in the AI era. 2 Literature review Previous studies on prompts for translation and LLMs have primarily been published in the field of natural language processing (NLP). Most research has fo89 Masaru Yamada cused on using LLMs as replacements for existing MT systems or as tools to address the shortcomings of MT. Specifically, these studies often investigate how translation accuracy compares between LLM-generated translations and traditional MT when using simple prompts, such as zero-shot prompts that merely instruct the model to translate from one language to another. For example, several studies (such as Hendy et al. 2023, Jiao et al. 2023, Wang et al. 2023) have examined how accurately LLMs can translate using zero-shot prompts and found that, for high-resource languages—languages with abundant digital resources and training data such as English, Spanish, or Chinese—LLMs can produce translations with accuracy comparable to or exceeding that of traditional MT (Way 2025). Other studies have explored using LLMs and prompts to enhance translation quality, particularly for low-resource languages. For instance, Jiao et al. (2023) analysed ChatGPT’s MT capabilities and found that while it competes with commercial systems for high-resource European languages, it struggles with lowresource and distant languages. Gao et al. (2023) developed a new method for translation prompts that includes task information, contextual domain information, and part-of-speech tags, which significantly improved ChatGPT’s performance, surpassing commercial systems in multiple translation directions. Zhang et al. (2023) provided a comprehensive summary of prompt strategies used to date and attempted to overcome the shortcomings in translation for low-resource language combinations and other challenging scenarios. Recent experiments have also focused on the translation of high-context and multi-modal materials such as Japanese manga using LLMs. Yang et al. (2024) conducted experiments demonstrating that applying appropriate prompts can significantly improve translation accuracy in such contexts. However, practical translation organisations have raised concerns about the feasibility of applying these methods to manga, particularly in capturing its nuanced cultural and contextual layers, which are integral to the genre (Japan Association of Translators 2024). While these studies originate from the NLP field and tend to focus on engineering improvements, there appears to be a difference in the understanding of translation between NLP researchers and Translation Studies (TS) scholars. In NLP, translation is often approached – particularly in evaluation – as if there were a single correct outcome. At the system level, however, MT models may generate multiple possible translations and then output the one judged to be the “most likely”, which does not necessarily align with what a human expert would consider the “best”. By contrast, TS generally recognises that translation may vary according to context and purpose. This variability makes it difficult 90 5 Teaching translation with AI to define what constitutes a “good” or “quality” translation. For instance, classical translation theories such as Skopos theory suggest that translation strategies (e.g. domestication vs. foreignisation (Venuti 1995), covert vs. overt translation (House 1981)) might depend on the purpose of the translation. Over time, TS has developed insights into these variations. The following sections attempt to address the gap left by NLP researchers by incorporating the metalanguages of TS, namely the concepts of translation, into LLM prompts and examining the resulting changes in translation output. Yamada (2023), for instance, investigated the impact of incorporating translation purposes and target audiences into prompts on ChatGPT’s translation output. This study focused on the pre-production phase of the translation process, drawing on previous translation research, industry practices, and ISO standards. By including concepts and terms from TS as prompts, the research explored the potential for achieving flexible translations that traditional MT systems have struggled to produce. The study evaluated changes in translation output using subjective qualitative assessments and cosine similarity, incorporating concepts such as dynamic equivalence (Nida & Taber 1969/2003). Further, He (2024) explored using translation research concepts to design prompts for LLMs to improve translation quality. The study discussed the effectiveness of incorporating conceptual tools and personas of translators and authors into ChatGPT’s translation task prompts. Although the small-scale experiments indicated limited effectiveness in improving translation quality, the paper highlighted the need for further research on the impact of TS concepts on LLMs translation tasks. Building on the ideas from these two papers, this chapter presents the study finding evaluating the effectiveness and potential of incorporating TS concepts as prompts for translator training, rather than assessing improvements in LLMgenerated translation quality. It further discusses the implications of this approach for translation education. 3 Aims and pedagogical potential of this chapter As observed in the comparative analysis of previous research, prompts in NLP often tend to prioritise the engineering aspect of achieving equivalence between source and target texts, typically measured by standardised translation metrics such as BLEU (Papineni et al. 2002) and COMET (Rei et al. 2020). This approach generally focuses on how closely the translation aligns with the source text. In contrast, TS often emphasises the importance of how the translation is received by the target audience, considering the context and situational variables. 91 Masaru Yamada This philosophical challenge in evaluating translation quality involves numerous variables and may not be solely based on source-target text equivalence. By incorporating the concepts and metalanguages of TS, it may be possible to create prompts for LLMs that better address these complex considerations. This chapter aims to serve as a starting point for designing pedagogical activities that leverage concepts and terminologies (metalanguages) from TS to create effective and unique prompts for LLMs. While not attempting to exhaustively evaluate all metalanguages as LLM prompts for educational use, the chapter presents several specific examples to inspire translation educators and instructors in their practice. By doing so, it highlights the potential of integrating TS knowledge into LLM-assisted translation education. To achieve this, the chapter first provides an overview of the foundational concepts of prompt engineering for LLMs, specifically few-shot prompts and CoT prompting. It then explores how these concepts might be integrated with existing TS knowledge. Through these examples and discussions, the chapter aims to offer practical suggestions and potential directions for incorporating TS concepts into LLM-assisted translation education, rather than presenting definitive solutions. 4 Few-Shot prompts and CoT prompting The concepts of few-shot prompts and CoT prompting are explained in the Prompt Engineering Guidelines.1Applying these concepts in this chapter requires some modifications and expanded interpretations, with concrete examples provided in Sections 5 and 6. Few-shot prompting is a prompt design method for LLMs that includes a few examples of input-output pairs along with task instructions. This approach allows the model to infer the task’s intent and output format from the provided examples, potentially leading to more accurate and desirable outputs. Few-shot prompting is particularly effective for complex tasks or when zero-shot prompting (instructions without examples) may be insufficient. Consider a scenario similar to a translation memory where portions of source and target texts are provided to the LLM. The model may learn implicit patterns from these examples and apply these learned patterns to generate translations. This method parallels actual practices in the translation industry, where translation service providers often provide translation memories to human translators. Translators typically learn from these past translations, observing how style guide rules are concretely applied. Similarly, translation coordinators or project 1https://www.promptingguide.ai 92 5 Teaching translation with AI managers provide instructions to human translators. This chapter aims to investigate how comparable instructions, when used as few-shot prompts for LLMs, might affect translation output. CoT prompting involves verbalising the thought process as a prompt. Just as humans follow a step-by-step reasoning process to solve problems, CoT prompting encourages LLMs to explicitly follow a CoT. This leads to more accurate and interpretable responses. In the context of translation, few-shot prompts provide examples of source and target texts without explaining the intermediate process. In contrast, CoT prompting could describe the translation process, referencing translation process research literature. For instance, a translator might read the source text, produce a literal translation, consider the context, monitor the translation, and then revise it to make sense for the target audience, culture, and readers (e.g., monitor model, Tirkkonen-Condit 2005). This step-by-step description of the translation process can be given as a CoT prompting. Additionally, as suggested by Yamada (2023), CoT prompting can be interpreted as detailed translation briefs. He (2024) demonstrated that setting translator personas (profiles) can also be considered a variant of CoT prompting. In professional translation settings, companies typically select suitable translators for specific tasks and provide detailed translation briefs. This chapter interprets CoT prompting as analogous to the language used for describing translation processes (Miyata et al. 2023) and investigates how such prompts might influence translation output, as well as their potential educational implications. 5 Example of a few-shot prompt Firstly, we present an example of a few-shot prompt. This example aims to examine how the output changes when a portion of a translation memory is input into the LLM. However, from an educational perspective, it is valuable to consider how learners, when faced with their own translation tasks, can study and emulate various linguistic aspects from the given corpus and incorporate them into their subsequent translations. In this case, we deliberately focused on linguistic tone and manner. A highly distinctive translation corpus was prepared. As indicated in the prompt below, the original English text was contextualised as a monologue reflection of an elderly person on their life. In the corresponding Japanese translation, we used expressions characterised by sentence endings typical of the linguistic style associated with Japanese elderly individuals, creating an exaggerated translation that exemplifies their manner of speaking. Although this example is quite ex93 Masaru Yamada treme and may lack realism, it can be likened to the specific style and voice used in a particular company’s user manuals in a real-world context. The purpose of this approach is twofold: to observe how the LLM adapts its output based on the provided translation memory, and to encourage learners to critically analyse how they can learn from and apply specific linguistic features, tones, and styles in their own translation work. This exercise, while using an exaggerated example, aims to heighten awareness of the importance of adapting language to specific contexts and audiences in translation practice. This prepared sample set was then given as a few-shot prompt to the LLM, instructing it to translate a new source text following the corpus. The LLM used was Claude 3.5 Sonnet. To provide a comparison, the new source text was also translated using the MT service DeepL. Subsequently, the LLM translations and the MT translations were compared against the original corpus. Additionally, evaluations were conducted using the automatic evaluation tool COMET2The results of these comparisons and evaluations are presented below. LLM Prompt 5.1: Below are provided the English text [source text], in which an elderly individual living in contemporary America reflects on their life, and its Japanese translation [target text]. Based on this corpus, please translate the [new source text] into Japanese. [source text] “When I was young, life was so different. We didn’t have all these gadgets and technology. We had to work hard for everything we had. But, those were good times too. I remember walking miles to school, playing outside until dusk, and everyone knew each other in the community.” [target text] 「若い頃はのう、生活はまったく違っとったんじゃ。今みたい なガジェットやテクノロジーはなかったんじゃ。何でも手に入れるために一生 懸命働かなければならんかったんじゃ。でも、それもまた良い時代だったんじ ゃよ。学校まで何マイルも歩いて、夕暮れまで外で遊んで、みんながコミュニ ティの中でお互いを知っとったんじゃ。」 2We used WMT20-COMET for this evaluation. 94 5 Teaching translation with AI [new source text] “I miss those simple days. We didn’t have much, but we were happy. I spent my summers helping my father on the farm and learning about life from him. Sundays were always special, with family gatherings and big dinners. As I grew older, the world changed rapidly, and sometimes it’s hard to keep up. But I cherish those memories, they keep me grounded.” Figure 1: Claude vs. DeepL. According to the COMET scores in Figure 1, Claude’s scores range from 0.360 to 0.995, while DeepL’s scores range from 0.633 to 1.051. These scores indicate generally high-quality translations, as scores closer to 1.0 typically reflect strong alignment with the reference text. On average, DeepL achieves higher scores at 0.626 compared to Claude at 0.499. However, when evaluated by a human translator, a stark stylistic difference becomes evident between the two. Claude Translation skilfully captures the distinctive elderly speech patterns found in the reference translations. Using a small corpus, it creates the impression that the same elderly person is continuing the conversation seamlessly. In contrast, DeepL Translation employs a completely standard Japanese tone, making it appear as though a different person is speaking, thereby disrupting the flow of the monologue. This discrepancy highlights a limitation of the COMET scoring system: it fails to account for stylistic features or cultural nuances such as those present in Japanese “elderly speech patterns”. 95 Masaru Yamada proach may offer opportunities to foster autonomous learning and develop essential qualities for independent translators. It is also important to acknowledge some limitations. Not all responses from LLMs were accurate, and some prompts were less successful than others. For instance, while the LLM correctly explained why COMET could not provide a fair evaluation in the few-shot prompt example, it failed to give a reasonable answer when asked which translation (Claude or DeepL) was closer to the reference translation. Such errors and limitations of LLMs become more apparent with increased use. However, given that translator training inherently requires maintaining a critical perspective, I believe exploring the possibilities of using LLMs is as important as considering the risks. In conclusion, this chapter has demonstrated concrete methods for exploring the potential of LLMs in translation education. By leveraging the concepts and metalanguages of TS in prompt engineering, we can create more engaging, interactive, and effective learning experiences for translation students. While challenges and limitations exist, the potential benefits of integrating LLMs into translator training are significant and warrant further investigation and development. References American Translators Association. 2024. ATA statement on artificial intelligence. https://www.atanet.org/advocacyoutreach/atastatementonartificialintelligence/. Dorst, Aletta G. 2024. Metaphor in literary machine translation: Style, creativity and literariness. In Andrew Rothwell, Andy Way & Roy Youdale (eds.), Computer-assisted literary translation, 173–186. New York, USA: Routledge. DOI: 10.4324/9781003357391-9. ELIS. 2025. ELIS 2025 European language industry survey. Tech. rep. ELIS Research. 1–53. http://elissurvey.org/wpcontent/uploads/2025/03/ELIS2025_Report.pdf. European Council of Literary Translators’ Associations. 2024. No one left behind, no language left behind, no book left behind. https://www.ceatl.eu/no-one-leftbehind-no-language-left-behind-no-book-left-behind. Gao, Yuan, Ruili Wang & Feng Hou. 2023. How to design translation prompts for ChatGPT: an empirical study. https://arxiv.org/abs/2304.02182. He, Sui. 2024. Prompting ChatGPT for translation: a comparative analysis of translation brief and persona prompts. https://arxiv.org/abs/2403.00127. 102 5 Teaching translation with AI Hendy, Amr, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify & Hany Hassan Awadalla. 2023. How good are GPT models at machine translation? A comprehensive evaluation.https://arxiv.org/abs/2302.09210. House, Juliane. 1981. A model for translation quality assessment. 2nd edn. Tübingen, Germany: Gunter Narr. ITI. 2024. Slow translation manifesto. https://www.iti.org.uk/discover/policy/ slow-translation-manifesto.html. Japan Association of Translators. 2024. Statement on the public and private sector initiative to use AI for high-volume translation and export of manga. https:// prtimes.jp/main/html/rd/p/000000001.000143535.html. Jiao, Wenxiang, Wenxuan Wang, Jen-tse Huang, Xing Wang, Shuming Shi & Zhaopeng Tu. 2023. Is ChatGPT a good translator? Yes with GPT-4 as the engine. https://arxiv.org/abs/2301.08745. LITHME Project. 2021. Language in the human-machine era. https://lithme.eu/. Miyata, Rei, Masaru Yamada & Kyo Kageura. 2023. Metalanguages for dissecting translation processes: Theoretical development and practical applications. London, UK: Routledge. https : / / www . routledge . com / Metalanguages - for - Dissecting-Translation-Processes-Theoretical-Development-and-PracticalApplications/Miyata-Yamada-Kageura/p/book/9781032168951. Nida, Eugene A. & Charles R. Taber. 1969/2003. The theory and practice of translation. Leiden, Netherlands: Brill. Papineni, Kishore, Salim Roukos, Todd Ward & Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, 311–318. Philadelphia, USA: Association for Computational Linguistics. DOI: 10.3115/ 1073083.1073135. https://www.aclweb.org/anthology/P02-1040. Rei, Ricardo, Craig Stewart, Ana C. Farinha & Alon Lavie. 2020. COMET: A neural framework for mt evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), 2685–2702. Online: Association for Computational Linguistics. DOI: 10.18653/v1/2020.emnlpmain.213. https://www.aclweb.org/anthology/2020.emnlp-main.213. Sayers, Dave, Rui Sousa-Silva, Sviatlana Höhn, Lule Ahmedi, Kais AllkiviMetsoja, Dimitra Anastasiou, Štefan Beňuš, Lynne Bowker, Eliot Bytyçi, Alejandro Catala, Anila Çepani, Rubén Chacón-Beltrán, Sami Dadi, Fisnik Dalipi, Vladimir Despotovic, Agnieszka Doczekalska, Sebastian Drude, Karën Fort, Robert Fuchs, Christian Galinski, Federico Gobbo, Tunga Gungor, Siwen Guo, Klaus Höckner, Petralea Láncos, Tomer Libal, Tommi Jantunen, Dewi Jones, Blanka Klimova, Eminerkan Korkmaz, Sepesy Maučec Mirjam, Miguel Melo, 103 Masaru Yamada Fanny Meunier, Bettina Migge, Barbu Mititelu Verginica, Aurélie Névéol, Arianna Rossi, Antonio Pareja-Lora, Christina Sanchez-Stockhammer, Aysel Şahin, Angela Soltan, Claudia Soria, Sarang Shaikh, Marco Turchi & Sule Yildirim Yayilgan. 2021. The dawn of the human-machine era: A forecast of new and emerging language technologies. https://hal.science/hal-03230287. Slator. 2024. French translators society takes tough stance on AI translation. https:// slator.com/french-translators-society-takes-tough-stance-on-ai-translationgenai/. Tirkkonen-Condit, Sonja. 2005. The monitor model revisited: Evidence from process research. Meta 50(2). 405–414. DOI: 10.7202/010990ar. Venuti, Lawrence. 1995. The translator’s invisibility: A history of translation. London, UK: Routledge. Vinay, Jean-Paul & Jean Darbelnet. 1958/1995. Comparative stylistics of French and English: A methodology for translation. Amsterdam, Netherlands: John Benjamins. DOI: 10.1075/btl.11. Wang, Longyue, Chenyang Lyu, Tianbo Ji, Zhirui Zhang, Dian Yu, Shuming Shi & Zhaopeng Tu. 2023. Document-level machine translation with large language models. In Houda Bouamor, Juan Pino & Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 16646– 16661. Singapore: Association for Computational Linguistics. DOI: 10.18653/v1/ 2023.emnlp-main.1036. Way, Andy. 2025. What does the future hold for translation technologies? In Stefan Baumgarten & Michael Tieber (eds.), The Routledge handbook of translation technology and society, 448–461. London, UK: Routledge. Yamada, Masaru. 2023. Optimizing machine translation through prompt engineering: An investigation into ChatGPT’s customizability. In Masaru Yamada & Félix do Carmo (eds.), Proceedings of Machine Translation Summit XIX, Vol. 2: Users Track, 195–204. Macau, China: Asia-Pacific Association for Machine Translation. https://aclanthology.org/2023.mtsummit-users.19. Yang, Zhishen, Tosho Hirasawa, Edison Marrese-Taylor & Naoaki Okazaki. 2024. Large language models as manga translators: A case study. In Proceedings of the 30th Annual Meeting of the Association for Natural Language Processing, 2012–2017. Tokyo, Japan: The Association for Natural Language Processing. https://www.anlp.jp/proceedings/annual_meeting/2024/pdf_dir/P7-13.pdf. Zhang, Biao, Barry Haddow & Alexandra Birch. 2023. Prompting large language model for machine translation: A case study. In Proceedings of the 40th International Conference on Machine Learning (ICML’23). Honolulu, USA: JMLR.org. 104