Prompt Design for Large Language Models in Supply Chain Risk Detection: A Case Study on Cambridge Taxonomy–Based News Classification
| dc.contributor.author | Ali, Rehan | |
| dc.contributor.department | fi=Tietotekniikan laitos|en=Department of Computing| | |
| dc.contributor.faculty | fi=Teknillinen tiedekunta|en=Faculty of Technology| | |
| dc.contributor.studysubject | fi=Tietotekniikka|en=Information and Communication Technology| | |
| dc.date.accessioned | 2026-08-03T19:31:37Z | |
| dc.date.issued | 2026-07-07 | |
| dc.description.abstract | In global e‑commerce, supply chain risk (SCR) is becoming a more common signal source. While this information is often identified in unstructured business news, it is hard to translate these signals into structured labels aligned with the taxonomy. Manual annotation is not scalable, and current large language model (LLM) methods may underestimate or overestimate risk classes or fail to capture nuances in geopolitical signals. This thesis investigates the impact of prompt design and retrieval strategy on Cambridge Taxonomy of Business Risks (CTBR)‑aligned SCR classification from news using state‑of‑the‑art LLMs. Using a case study of 121 CTBR‑labelled news articles about major steel producers, two instruction‑tuned LLMs (LLaMA‑3‑70B and GPT‑4o) are evaluated under three prompt structures: Simple, Strict and Ontology‑guided. In all experiments, a fixed few‑shot semantic retrieval setup and a Leave‑One‑Company‑and‑Year‑Out cross‑validation scheme are used to ensure that the only differences are the prompt structure and the model family. This study focuses on over- and under-prediction, the impact of class imbalance on recall, and Geopolitical versus Governance and Financial risks. Empirical results show that LLaMA‑3‑70B with Strict prompt achieves the best overall performance, with GPT‑4o being more sensitive to the impact of alignment, significantly over‑using the No Risk label unless provided with an ontology‑style prompt. This holds true for rare classes like Technology and Environmental, where retrieval sparsity is a problem for both models. Refining the ontology prompt \texttt{Ontology\_v2} in an exploratory manner results in substantial gains in GPT‑4o's class‑level F1 scores, and the number of No Risk false positives drops drastically. The thesis presents a systematic comparison of prompt variants for taxonomy‑aligned SCR detection and practical guidelines for prompt and retrieval design in LLM‑based risk monitoring pipelines. It concludes that fully autonomous risk classification is currently not possible, but that few-shot, carefully engineered prompts, coupled with semantic retrieval and ontology guidance, are an effective way to develop more reliable and reusable SCR monitoring systems. | |
| dc.format.extent | 100 | |
| dc.identifier.uri | https://www.utupub.fi/handle/11111/62860 | |
| dc.identifier.urn | URN:NBN:fi-fe20260803114508 | |
| dc.language.iso | eng | |
| dc.rights | fi=Julkaisu on tekijänoikeussäännösten alainen. Teosta voi lukea ja tulostaa henkilökohtaista käyttöä varten. Käyttö kaupallisiin tarkoituksiin on kielletty.|en=This publication is copyrighted. You may download, display and print it for Your own personal use. Commercial use is prohibited.| | |
| dc.rights.accessrights | suljettu | |
| dc.subject | supply chain risk | |
| dc.subject | Cambridge Risk Taxonomy | |
| dc.subject | large language models | |
| dc.subject | prompt engineering | |
| dc.subject | few‑shot retrieval | |
| dc.subject | LLaMA‑3‑70B | |
| dc.subject | and GPT‑4o | |
| dc.title | Prompt Design for Large Language Models in Supply Chain Risk Detection: A Case Study on Cambridge Taxonomy–Based News Classification | |
| dc.type.ontasot | fi=Diplomityö|en=Master's thesis| |
Tiedostot
1 - 1 / 1