The Illusion of Private AI: Rethinking Data Privacy in Isolated Contextand Private Knowledge Base Systems
DOI:
https://doi.org/10.64229/b98f3723Keywords:
Generative AI, Private AI, NotebookLM, Data privacy, Large language models, AI security, Retrieval-augmented generation, Governance, Zero trust, Higher educationAbstract
Private AI is used in this paper as a deployment-based umbrella term for AI systems whose data and computation are constrained within a defined trust boundary. That boundary may be local or on-device, on-premises or private-network infrastructure, a private or dedicated cloud, or a managed cloud service that provides logical and contractual isolation. These models are not equivalent. This paper is concerned mainly with the final category: cloud-based isolated-context AI services that allow users to upload a bounded collection of documents and query them through one or more foundation models. NotebookLM is used as the main illustrative case, while Microsoft 365 Copilot and ChatGPT Business/Enterprise are comparison points. The central argument is straightforward: source isolation and exclusion from foundation-model training are useful privacy controls, yet neither is equivalent to local processing. The study uses a structured conceptual review of peer-reviewed security and privacy research, authoritative standards and regulator guidance, and current provider documentation. It separates privacy, cybersecurity and governance concerns across five layers: transmission and account boundary; inference; stored and derived representations; training and model improvement; and governance and human access. The paper then converts those layers into a practical zero-trust decision framework for deciding what may be uploaded, what requires an approved enterprise environment, and what should be minimised, anonymised or kept local. Its practical value is that it gives educators, managers and technical staff a way to make a data decision without pretending that a simple label such as private or no training answers every privacy question. The framework is illustrated through the design of an AI Security, Ethics and Management module at the University of Leicester. No empirical validation is claimed at this phase; the educational application is a design case that motivates a planned evaluation.
References
[1]Stephenson R, Armstrong C. Student generative artificial intelligence survey 2026. London: Higher Education Policy Institute. 2026. (HEPI Report 199). Available from: https://www.hepi.ac.uk/reports/student-generative-ai-survey-2026/ (accessed on 07 September 2026).
[2]Microsoft, LinkedIn. 2024 Work trend index annual report: AI at work is here. Now Comes the Hard Part. 2024. Available from: https://www.microsoft.com/en-us/worklab/work-trend-index/ai-at-work-is-here-now-comes-the-hard-part/ (accessed on 07 September 2026).
[3]Google. Privacy and terms of use in NotebookLM. NotebookLM Help. Available from: https://support.google.com/notebooklm/answer/17004255 (accessed on 07 September 2026).
[4]Google. Learn about NotebookLM. NotebookLM Help. Available from: https://support.google.com/notebooklm/answer/16164461 (accessed on 07 September 2026).
[5]Google Workspace. Generative AI in Google Workspace privacy hub. Google Workspace Help. Available from: https://knowledge.workspace.google.com/admin/gemini/generative-ai-in-google-workspace-privacy-hub?hl=en (accessed on 07 September 2026).
[6]National Institute of Standards and Technology. Artificial intelligence risk management framework (AI RMF 1.0), NIST AI 100-1. 2023. DOI: 10.6028/NIST.AI.100-1
[7]National Institute of Standards and Technology. Artificial intelligence risk management framework: Generative Artificial Intelligence Profile, NIST AI 600-1. 2024. DOI: 10.6028/NIST.AI.600-1
[8]Information Commissioner’s Office. Guidance on AI and data protection. 2023. Available from: https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/ (accessed on 07 September 2026).
[9]European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence. 2024. Available from: https://eur-lex.europa.eu/eli/reg/2024/1689/oj (accessed on 07 September 2026).
[10]Yao Y, Duan J, Xu K, Cai Y, Sun Z, Zhang Y. A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly. High-Confidence Computing, 2024, 4(2), 100211. DOI: 10.1016/j.hcc.2024.100211.
[11]Kibriya H, Khan WZ, Siddiqa A, Khan MK. Privacy issues in Large Language Models: A survey. Computers and Electrical Engineering, 2024, 120, 109698. DOI: 10.1016/j.compeleceng.2024.109698.
[12]Zhang R, Li HW, Qian XY, Jiang WB, Chen HX. On large language models safety, security, and privacy: A survey. Journal of Electronic Science and Technology, 2025, 23(1), 100301. DOI: 10.1016/j.jnlest.2025.100301
[13]Microsoft. Enterprise data protection in Microsoft Copilot and Microsoft Copilot Chat. Microsoft Learn. Available from: https://learn.microsoft.com/en-us/microsoft-365/copilot/enterprise-data-protection (accessed on 07 September 2026).
[14]Microsoft. Frequently asked questions about Microsoft Copilot Chat. Microsoft Learn. Available from: https://learn.microsoft.com/en-us/copilot/faq (accessed on 07 September 2026).
[15]OpenAI. Business data privacy, security, and compliance. 2026. Available from: https://openai.com/business-data/ (accessed on 07 September 2026).
[16]Microsoft. Microsoft 365 Copilot architecture and how it works. Microsoft Learn. Available from: https://learn.microsoft.com/en-us/microsoft-365/copilot/microsoft-365-copilot-architecture (accessed on 07 September 2026).
[17]OpenAI. Company knowledge in ChatGPT (Business, Enterprise, and Edu). OpenAI Help Center. https://help.openai.com/en/articles/12628342-company-knowledge-in-chatgpt-business-enterprise-and-edu (accessed on 07 September 2026).
[18]Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 2020, 33. DOI: 10.48550/arXiv.2005.11401
[19]Zeng S, Zhang J, He P, Xing Y, Liu Y, Xu H, et al. The good and the bad: Exploring privacy issues in retrieval-augmented generation (RAG). In: Findings of the Association for Computational Linguistics: ACL 2024. Bangkok, Thailand: Association for Computational Linguistics, 2024, 4505-4524. DOI: 10.18653/v1/2024.findings-acl.267
[20]Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. In: Advances in Neural Information Processing Systems 30 (NeurIPS 2017). Long Beach, CA, USA: Curran Associates Inc. 2017, 5998-6008. DOI: 10.48550/arXiv.1706.03762.
[21]Akyürek E, Schuurmans D, Andreas J, Ma T, Zhou D. What learning algorithm is in-context learning? Investigations with linear models. arXiv, 2022 DOI: 10.48550/arXiv.2211.15661
[22]von Oswald J, Niklasson E, Randazzo E, Sacramento J, Mordvintsev A, Zhmoginov A, et al. Transformers learn in-context by gradient descent. In: Proceedings of the 40th International Conference on Machine Learning, PMLR 202, 2023. Available from: https://proceedings.mlr.press/v202/von-oswald23a.html (accessed on 07 September 2026).
[23]Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, et al. Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems, 2020. DOI: 10.48550/arXiv.2005.14165
[24]Carlini N, Tramèr F, Wallace E, Jagielski M, Herbert-Voss A, Lee K, et al. Extracting training data from large language models. In: 30th USENIX Security Symposium, 2021. Available from: https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting (accessed on 07 September 2026).
[25]Morris J, Kuleshov V, Shmatikov V, Rush A. Text embeddings reveal (Almost) as much as text. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Singapore: Association for Computational Linguistics, 2023, 12448-12460. DOI: 10.18653/v1/2023.emnlp-main.765
[26]Zyskind G, South T, Pentland A. Don’t forget private retrieval: Distributed private similarity search for large language models. Proceedings of the Fifth Workshop on Privacy in Natural Language Processing, 2024, 7-19. Available from: https://aclanthology.org/2024.privatenlp-1.2/ (accessed on 07 September 2026).
[27]Greshake K, Abdelnabi S, Mishra S, Endres C, Holz T, Fritz M. Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, 2023. DOI: 10.1145/3605764.3623985
[28]Tan Z, Zhao C, Moraffah R, Li Y, Wang S, Li J, et al. Glue pizza and eat rocks–exploiting vulnerabilities in retrieval-augmented generative models. Proceedings of EMNLP 2024, 1610-1626. DOI: 10.18653/v1/2024.emnlp-main.96
[29]Debenedetti E, Severi G, Carlini N, Choquette-Choo CA, Jagielski M, Nasr M, et al. Privacy side channels in machine learning systems. 33rd USENIX Security Symposium, 2024. Available from: https://www.usenix.org/conference/usenixsecurity24/presentation/debenedetti (accessed on 07 September 2026).
[30]OWASP GenAI Security Project. OWASP Top 10 for LLM Applications 2025. Available from: https://genai.owasp.org/llm-top-10/ (accessed on 07 September 2026).
[31]OWASP GenAI Security Project. LLM08: 2025 Vector and Embedding Weaknesses. 2025. Available from: https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/ (accessed on 07 September 2026).
[32]UK National Cyber Security Centre. Prompt injection is not SQL injection (it may be worse). 2025. Available from: https://www.ncsc.gov.uk/blog-post/prompt-injection-is-not-sql-injection (accessed on 07 September 2026).
[33]UK National Cyber Security Centre. Thinking about the security of AI systems. 2023. Available from: https://www.ncsc.gov.uk/blog-post/thinking-about-security-ai-systems (accessed on 07 September 2026).
[34]Bommasani R, Hudson DA, Adeli E, Altman R, Arora S, von Arx S, et al. On the opportunities and risks of foundation models. 2021. DOI: 10.48550/arXiv.2108.07258
[35]Weidinger L, Mellor J, Rauh M, Griffin C, Uesato J, Huang PS, et al. Ethical and social risks of harm from Language Models. arXiv, 2021. DOI: 10.48550/arXiv.2112.04359
[36]International Organization for Standardization. ISO/IEC 42001:2023, Information technology–Artificial intelligence–Management system, 2023. Available from: https://www.iso.org/standard/42001 (accessed on 07 September 2026).
[37]International Organization for Standardization. ISO/IEC 23894:2023, Information technology–Artificial intelligence–Guidance on risk management, 2023. Available from: https://www.iso.org/standard/77304.html (accessed on 07 September 2026).
[38]Russell SJ, Norvig P. Artificial intelligence: A modern approach, 4th ed. NJ: Pearson, 2021. Available from: https://www.pearson.com/en-us/subject-catalog/p/artificial-intelligence-a-modern-approach/P200000003500/9780137505135 (accessed on 07 September 2026).
[39]Goodfellow I, Bengio Y, Courville A. Deep Learning. MIT Press, 2016. Available from: https://www.deeplearningbook.org/ (accessed on 07 September 2026).
[40]Mualla KJ, Mualla W. Leveraging AI in Higher Education: Contemporary Approaches for Teaching, Learning and Assessment Design. In: Higher Education in the Arab World: Artificial Intelligence, Springer, Cham, 2026, 169-191. DOI: 10.1007/978-3-031-99068-7_8
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Karim Mualla (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.