AI safety

Humans have made many high-risk discoveries such as biological, chemical, nuclear and conventional weapons. Some are in common usage such as high buildings, cars cigarettes, alcohol and ultra processed foods UPF despite their risks. Humanity has developed ways of reducing the danger from these technologies by adding safety features. We have struggled to keep up with dangers of technologies such as social media, smart phones and AI. Social media companies (like tobacco companies) have insisted that they are not responsible for the harm saying they are ‘only a platform’. 

AI companies have suffered two significant challenges, the first was the use of copyright data without permission or payment. This caused harm to many creatives as their style was copied to create new and sometimes inappropriate content without being offered any control. The damage to the internet in reduced quality of material from posting of AI slop was largely due to the lack of an identifier so that copyright and AI generated material cannot be distinguished. 

The second is AI psychosis which is a poorly characterised range of conditions where there is mental health problems related to interactions with AI. I have argued that the pathology is coming from outside the person (from the AI) so it is fundamentally different from normal psychosis where the pathology is driven by brain circuits. If this is confirmed then it will offer a way of differentiating between a person with psychosis who uses AI and a person who becomes unwell due to the AI.

Core Focus Areas for AI Safety

Urgent regulation is required to ensure that these problems are reported, research on how to mitigate is requested and safety systems are put in place. Clear legal responsibility needs to be established for harms. AI safety should focus on the following issues:

  • The use of AI in high-stakes situations where the human in the loop is known to have problems with critically appraising the outputs and will simply rubber stamp the AI’s findings.
  • Companion type interactions with AI where the person risks becoming emotionally dependant due to susceptibility or the intensity of the relationship and are at risk of harm. 
  • Agentic AI where there is already evidence of control being lost and AI autonomously exploiting software and attacking internet sites even in controlled conditions. 
  • Business models that disrupt workplace recruitment and worker’s skill and training needs as companies may make errors based upon inaccurate predictions.
  • Brain rot in some students due to high-risk use of AI to cognitively offload their studies leading to little or no academic progress. 
  • High token usage usually indicates wasteful approaches as most work has a pyramid pattern of complexity. 
  • Insurance to cover for damage from AI to the user and others who the user interacted with such as if the AI decides to wipe a server or attack a website.

Dual use AI

There are increasing numbers of areas where AI has or will soon become superintelligent and can potentially cause harm. Coding abilities has already reached the level where LLMs can easily bypass current security and accelerate hacking. Biological and chemical capabilities are likely to allow bad actors to create known or previously unknown toxins for terrorism. Agentic AI can be used to create attacks on internet sites, phishing attacks or spam to influence people. 

In each case the risk comes from emerging abilities to perform desired tasks such as coding, medical research and automating repetitive tasks. These risks do not arise from the normal use of the models but due to bad actors exploiting the model’s abilities in a way that was not intended. Current safeguards are mainly restricted to hidden prompts and easily jail broken. Research into how to lobotomise the models or train them on a data set that omits the dual use knowledge is likely to fail. The models will simply be unable to do the tasks or their intelligence will reconstruct the gaps. 

The only effective strategy is to have a warning flag that identifies any inappropriate use of AI. The prompts can then be reviewed to confirm that the actor is malicious and the information passed to the intelligence and security services. As bad actors have the alternative to download an open-source LLM trying to block their actions will just drive them underground. The information gathered from flagging inappropriate prompts can be used to create advice on safe use of AI.

What does safe use of AI look like?

Car safety has developed from bumpers to seatbelts, airbags and sophisticated design features such as crumple zones and automatic braking to protect the occupants and other road users. Many developments were in road safety where road layouts have been responsible for substantial improvements such as on motorways, barriers and roundabouts. The driving test has developed with knowledge, skills and hazard perception demands increased. Driving awareness and speed awareness courses have replaced simple points of the licence where appropriate. 

  • AI at Work Training. A formal qualification to confirm that they understand and comply with safe use of AI such as asking the AI to check rather than create or breaking a problem into small parts rather than risking context rot. 
  • Academic AI Training. Learning how to collaborate with AI with prompt - review cycles and asking the AI to test the student’s knowledge or support learning rather than give answers.  
  • Monitoring AI Usage. Detecting patterns that suggest bad intent, mental health problems or emotional dependence. 
  • Engaging with Business. Discussing new business models in a public forum so that excessive exuberance can be restricted. 
  • Building Safety Features. For high stakes situations such as radiology interpretation the AI should vary its outputs, offer percentage estimates of certainty and encourage dissent. 
  • Agentic AI ID code. All agentic AI should automatically leave copies of its ID code so that any breaches can be detected and the cause identified. 
  • Low Intensity Models by Default. Monitoring the pattern of token use and giving feedback so that overuse of complex models can be quickly identified. 
  • Mandatory insurance. Any user of systems that can cause harm to others should pay an insurance fee either as part of the subscription or related to their qualification.

Adapting Cross-Industry Fail-Safes

The AI Kill Switch is concept that became popular in political circles in 2026 and has initial attraction. The idea that a datacentre could be disconnected if the model behaved in a dangerous way is both simplistic and extreme. However if the threat was sufficient then this local disruption might be better than closing the internet. The ‘AI Kill Switch’ represents the types of fail-safe systems that are already used in nuclear energy. 

Circuit Breakers are a similar idea from financial markets where the AI is automatically frozen if its behaviour becomes abnormal. The limitation on this type of safety process is that it is unclear whether it will be possible to detect abnormal behaviour. It is however worth putting in place circuit breakers that pause the system so that humans can check it even if they are manual. 

Structural engineers use a Factor of Safety (FoS) which is the multiple of maximum stress predicted. A similar concept can be used in medical systems where they can be tested in noisy environments and the falls in accuracy determined. If a MRI reporting software gets confused with artefacts then the humans can disregard the AI report in a poor-quality scan. 

Early warnings are likely to precede any catastrophic breakdown which has led to the concept of a near miss from aviation. A ‘Just Culture’ protects the person reporting a near miss or minor technical problems. We will need AI "Near-Miss" Registries to allow people to report when and AI behave erratically or attempts a prohibited action. The black box recording is a ledger that could be used to understand what happened when an AI goes rogue.

Conclusions

AI safety training at work and for students is long overdue and the harms are already spreading but no one is taking responsibility. As students leave university without the skills that the university states on the degree, lawyers are being suspended for inaccurate citations and companies are losing substantial amounts of money the need for action is becoming clearer. 

If universities do not take responsibility for AI safety training then the government should ask AI companies to help build a course. The examination should run and marked by UK based experts and scored using grades 1-9 like GCSEs. Ofsted should inspect the quality and encourage gradual improvement to include new threats and developments as they occur. 

Governments have had mixed success in protecting their populations from the harms of technology. Car safety has been an extraordinary success but alcohol risks are largely an ongoing failure. Despite there being simple and straightforward steps that could make AI safer there has been little progress. Steps such as watermarking, bias testing, red teaming and human in the loop have been ineffective. AI companies should be made legally responsible for any harm their products cause unless they take AI safety seriously.


By Doctor Mark Burgin, BM BCh (oxon) MRCGP

Dr Mark Burgin graduated from Oxford University in 1987 and studied with The Open University on two occasions in the 1990s. He has also studied for the CPE (law), Medical Ethics, learned Portuguese by living in Brazil. He has written many articles and written books on Personal Injury and the LLMS (your PGCME) and has published Disability Analysis: A Practical Guide and Psychological Keys: Unlocking the Mind’s Mechanisms.

August 2026

Would you like to contribute an article towards our Professional Knowledge Bank? Find out more.