Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 184
Back to Search View Original Cite This Article

Abstract

<jats:p>Generative artificial intelligence (AI) coding assistants can substantially accelerate research workflows in ecology, fisheries science, and related quantitative disciplines, but they also create new pathways for accidental or adversarial disclosure of confidential information. Researchers in these fields routinely work with legally protected and commercially sensitive records (e.g., individual-level records, accurate geographic locations, pre-release data), and even small snippets of code context, console output, comments, or file paths can be disclosive when transmitted to cloud-hosted large language model (LLM) services. This paper provides a practical guide for researchers and institutions to use LLM-based coding assistants while maintaining confidentiality. We (i) catalogue common AI integration points in widely used development environments and R tooling, highlighting features that may activate or transmit context without obvious user intent; (ii) summarise key data exposure pathways – intentional sharing, accidental transmission, and adversarial exploitation such as prompt injection; (iii) propose a four-scenario framework spanning browser chat, editor autocomplete, agentic assistants, and local models, with increasing capability and risk; and (iv) offer actionable mitigations centered on a “develop on simulated data, run on real data” workflow and a complementary two-computer separation strategy. We also introduce the confideR R package, which supports session auditing, risk reminders, script scanning for sensitive text, and generation of prompt-ready data fingerprints and simulated datasets. Commercial tools for running LLMs safely are likely to rapidly improve in usability and capability, but so is the sophistication of malicious attacks. Our recommended approach is robust to future changes, because it uses simulated data and the physical separation of confidential data and LLM tools. Together, these recommendations aim to enable responsible uptake of AI coding assistants in the environmental sciences without compromising data confidentiality.</jats:p>

Show More

Keywords

data assistants coding simulated also

Related Articles


Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 76
PORE

About

Connect