In April 2026, Bee Cheng Hiang accidentally exposed the email addresses of 95,364 customers after an employee used a generative AI tool to create Python code for a marketing email campaign.
The incident became the first AI-related data breach notified to Singapore's Personal Data Protection Commission (PDPC). Marketing emails were sent in batches of around 1,000 customers, with recipients able to see other customers' email addresses within the same batch.
It would be easy to reduce the incident to "AI-generated code caused a data breach", but that misses the more useful engineering lesson.
The PDPC said the AI tool itself did not malfunction. The employee's prompt did not specify that each recipient's email address should be hidden from the others. The generated code was then deployed without sufficiently robust testing or independent supervisory review.
For businesses starting to use AI-assisted coding, that distinction matters.
Code can execute successfully and still do something completely different from what you intended.
What Happened?
A Bee Cheng Hiang employee used a generative AI tool to write a Python program for distributing a marketing email using a local customer list.
The employee instructed the tool to send the emails in batches but did not specify that the addresses of individual recipients should be hidden from one another.
The resulting implementation placed multiple customer email addresses into the same recipient field.
According to the PDPC's explanation reported by The Straits Times, the visual difference between the intended code and problematic code came down to the placement of a couple of brackets. That small change altered how the recipients were handled.
The emails were sent on April 25, 2026. The PDPC was notified on April 27.
Only customer email addresses were affected, and the PDPC said there was no evidence that the exposed information had subsequently been misused.
After discovering the problem, Bee Cheng Hiang stopped the bulk email process, corrected the script and informed affected customers.
The company also introduced a requirement for at least two staff members to verify bulk email communications before they are sent.
The Code Ran Successfully
One of the easiest misconceptions in software development is:
The code ran without errors, therefore it works.
But those are two different things...
Imagine a program designed to calculate a 9% charge but accidentally applying 90%.
The application might run normally. The server does not crash. There may be no error message in the logs. However, the calculation is still wrong.
This same principle applies to the Bee Cheng Hiang incident.
The script was capable of sending emails. The failure was in how those emails were addressed.
This becomes increasingly important as employees use AI tools to generate scripts without necessarily being experienced developers.
Running the program is only one test.
Someone else still has to check whether the result matches what the business actually intended.
The Testing Missed the Actual Output
According to the PDPC, testing was performed by checking activity logs without reviewing the contents of the actual test email.
For this particular system, the received email was the output that mattered.
A log could confirm that an email was generated and sent successfully. It would not necessarily reveal that hundreds of other customer addresses appeared in the recipient field.
A more realistic test could have used several controlled dummy email accounts.
Someone could then open each message and check:
- Whether the correct recipient received it
- Whether other recipients were visible
- Whether personalised information belonged to the correct customer
- Whether the subject and content were correct
- Whether links and attachments worked properly
The PDPC said Bee Cheng Hiang's follow-up measures would include testing emails using dummy accounts before deployment.
This principle applies well beyond email systems.
If software produces something a customer will eventually see, test the actual output.
AI-Generated Code Still Needs Code Review
AI coding tools are extremely good at turning a written instruction into working code quickly.
The problem is that they have to work with the requirement they were given.
Consider a prompt such as:
Write a Python script that sends this marketing email to my customer list in batches of 1,000.
There are unanswered questions inside that request.
- Should recipients see one another?
- Is everybody receiving identical content?
- Is customer information personalised?
- What happens if a batch fails halfway?
- Can the script send the same batch twice?
- Can the operation be stopped after it begins?
An experienced developer reviewing the requirement would probably think about at least some of these cases before deploying it.
Someone using an AI tool may not know that these questions need to be asked. They assume that generative AI would do the due diligence for them when they give it a prompt.
The generated code can still look completely reasonable.
And... this is where the code review matters.
AI lowers the amount of technical knowledge required to create code. It does not automatically give the person generating that code the experience required to identify everything that could go wrong.
Personal Data Raises the Stakes
Not every script needs the same review process.
A small internal program that renames ten test files has a very different risk profile from code capable of processing 95,000 customer records.
Businesses should consider what the code can affect.
Extra review makes sense when software can:
- Access personal information
- Send customer communications in bulk
- Delete or modify large amounts of data
- Process payments
- Change production databases
- Trigger transactions in another system
The PDPC advised organisations to conduct appropriate data protection impact assessments before adopting AI tools for business processes involving personal data.
It also recommended policies, processes, testing and review mechanisms around employees' use of AI tools.
A small company does not necessarily need a complicated AI governance department, but it does need to know when a script has enough access to cause serious damage if something goes wrong.
Put Safeguards Around Bulk Actions
There is another lesson here that applies whether the code was written by AI or a human developer.
Large actions deserve safeguards.
If an email system suddenly attempts to place hundreds of customer addresses into one recipient field, software can potentially detect that before anything is sent.
Bee Cheng Hiang committed to implementing automated measures intended to block mass email distribution when multiple email addresses appear within a single email field.
This is good, and the same idea can be used elsewhere.
A system could require confirmation before deleting 10,000 records.
A large refund could require another employee's approval, also known as "Dual Control Verification".
A bulk messaging system could show how many recipients are about to receive a message before allowing the send.
A data migration might process a small sample first rather than immediately changing the entire production database.
These controls do not make software infallible. Although they reduce the chance that one mistake immediately turns into a large incident, which is better than no safeguards at all.
Review the Output, Not Only the Code
Technical code review is useful, but businesses should also test what the software actually produces.
This is particularly valuable when the person using AI is not an experienced programmer.
They may struggle to recognise a subtle problem inside a Python script.
They can still inspect the outcome.
For an email script: Open the email.
For an invoice generator: Compare the generated invoice against the original information.
For a document automation system: Inspect the generated document.
For a data importer: Start with a small test dataset and inspect the records created.
For a customer portal: Log in as different types of users and check what each account can actually access.
You do not need to understand every line of code to notice that the output is wrong.
Imagine AI writing code that does not align with your intended outcome, but you continue building ontop of it. What happens? You get technical debt. Problems that cause you more trouble and time fixing it down the road.
Production Should Not Be the Test Environment
AI makes it possible to move from an idea to runnable code extremely quickly.
That speed can make it tempting to generate a script, run it against live data and see what happens.
Anything affecting customers or important business information deserves a safer process.
The exact setup depends on what is being developed.
For a simple mailing script, a controlled list of dummy addresses might be enough.
For a larger application, we would normally use a staging environment where changes can be tested without affecting the production system.
The higher the consequence of failure, the stronger that separation should be.
Our web application development process goes further into how staging, testing and deployment fit into a normal development workflow.
AI Changes Who Can Write Software
Generative AI has made software development accessible to many more people.
Someone who does not know Python can now describe a process in plain English and receive a functioning script seconds later.
Small companies can automate tasks that might previously have remained manual because bringing in a developer for every small script was impractical.
But with that, comes a new risk.
Someone can now produce software capable of accessing thousands of customer records without fully understanding what the generated code is doing.
If the employee does not know that bulk email recipients need to be handled separately, they may not know that this should be included in the prompt.
They may also be unable to identify the issue by looking at the resulting code.
The ability to generate software has become easier but the knowledge required to validate that software has not disappeared.
Should Companies Ban AI-Generated Code?
For most businesses, banning AI-assisted coding entirely would probably throw away useful technology rather than solve the underlying problem.
Developers can use AI tools to speed up repetitive tasks, understand unfamiliar code and create initial implementations faster.
Non-developers can use them for smaller internal automations.
The controls should match the risk of the task.
Generating a small script against disposable test data is very different from allowing generated code to access a production customer database.
A company could require additional technical review when AI-generated code:
- Handles personal data
- Operates against production systems
- Sends external communications
- Processes payments
- Performs bulk operations
Bee Cheng Hiang's voluntary undertaking includes establishing a framework governing employee use of AI for coding and independent technical review of AI-generated code involving personal data.
That is a much more practical approach than assuming every piece of generated code carries the same level of risk.
A Practical Process for SMEs
An SME does not need the software governance process of a bank before somebody can use an AI-generated script.
For code interacting with customer or important business data, a reasonable workflow could be:
1. Generate or write the code.
2. Clearly define what the code is supposed to do.
3. Have somebody else review the requirement and implementation where the risk justifies it.
4. Run it against test data.
5. Inspect the actual output.
6. Test obvious failure cases.
7. Start with a small production batch where possible.
8. Monitor the first live run before allowing the process to continue at full scale.
The level of review should increase with the potential business impact.
A script transforming ten internal spreadsheet rows is not equivalent to a program about to contact 100,000 customers.
The Prompt Was Only Part of the Problem
The incident has understandably attracted attention because an incomplete AI prompt contributed to what happened.
Prompt quality is important.
It should not become the only lesson businesses take from the case.
Even a detailed prompt can still result in incorrect code.
A model can misunderstand an instruction, make an assumption the user did not anticipate, produce an incorrect implementation or fail to account for an edge case.
Better prompting reduces some risks.
Testing and review deal with the fact that those risks still exist.
A useful way to think about the prompt is as a set of requirements given to the AI.
Software projects have dealt with incomplete requirements long before generative AI existed.
AI Does Not Remove Responsibility for the Result
The PDPC specifically said this incident resulted from human error in developing the email distribution code with an AI tool and was not caused by a malfunction of the AI tool itself.
That is an important distinction for businesses introducing AI into their development processes.
Once AI-generated code becomes part of a production system, somebody still has to own the outcome and take responsibility for anything that goes wrong, especially code that affects other innocent users.
The tool can help write the implementation but the organisation still decides when that implementation is ready to touch real customers.
Frequently Asked Questions
How many Bee Cheng Hiang customers were affected?
95,364 customers had their email addresses exposed to other recipients of the marketing messages.
The PDPC said email addresses were the only personal data affected and that there was no evidence of further misuse.
Was Bee Cheng Hiang hacked?
No hacking was reported as the cause of this incident.
The disclosure resulted from how the generated bulk-email script handled recipient addresses.
Did the AI tool malfunction?
No.
The PDPC said the incident was not caused by an AI tool malfunction. The employee's prompt and the subsequent development, testing and review process led to the problem.
Was payment information exposed?
No payment information was reported as affected.
The data involved consisted of customer email addresses.
Can businesses safely use AI-generated code?
Yes, provided the code receives a level of testing and review appropriate to what it can affect.
Anything involving personal information, financial transactions, production databases or large customer-facing actions deserves greater scrutiny.
What did Bee Cheng Hiang change after the incident?
The company introduced double-verification for bulk emails and committed to measures including an AI coding governance framework, independent technical review of relevant AI-generated code, testing with dummy accounts and automated controls around bulk-email recipient fields.
Using AI-Generated Code in Your Business? Review It Before Production
AI can make a useful internal tool or automation much faster to build, but production code still needs to be treated as production code.
If your business is experimenting with AI-generated scripts, automations or changes to an existing web application, retroXpect can review the technical implementation, test how it behaves against the intended workflow and help identify issues before the code is allowed to interact with customers or live business data.
This can range from reviewing a small automation to taking over, fixing or extending an existing application.
If you have already built something using AI and are unsure whether it is ready for production, discuss the implementation with retroXpect before connecting it to live customer data.


