In April, Stanford University published the first in-depth study on AI psychosis, conducted with help from the Human Line Project. It studied the transcripts of 19 human/chatbot interactions…
OpenAI and Google helped fund the study.
The study itself says: “This work was supported by API credit grants from OpenAI and Google (Gemini) as well as from a gift from OpenAI.“
This wasn’t the first time Stanford and OpenAI have worked together on this topic, but it was the first time they had organized participation from victims who were harmed. Did those victims know who was funding that study? Would they have consented to participation if they knew the same company that caused their harm was financing it?
The work between OpenAI and Stanford to investigate harms of LLMs goes back to at least 2020.
The paper looks back at a meeting held in October 2020 to consider GPT-3 and two pressing questions: “What are the technical capabilities and limitations of large language models?” and “What are the societal effects of widespread use of large language models?” Coauthors of the paper described “a sense of urgency to make progress sooner than later in answering these questions.”
The record clearly shows that in 2020, OpenAI expected widespread harmful societal effects and barreled forward anyway.
Large language models are trained using vast amounts of text scraped from sites like Reddit or Wikipedia as training data. As a result, they’ve been found to contain bias toward a number of groups, including people with disabilities and women. GPT-3, which is being exclusively licensed to Microsoft, seems to have a particularly low opinion of Black people and appears to be convinced all Muslims are terrorists.
Who thought it was a good idea to train a new INTELLIGENCE system on the least intelligent and most emotionally immature text that has ever been made available – anonymous internet comments? If there is one thing I would encourage a developing human not to read and emulate – it would be anonymous online comments.
In a presentation about experiments probing anti-Muslim bias in GPT-3 presented at the first Muslims in AI workshop at NeurIPS, Abid described anti-Muslim bias demonstrated by GPT-3 as persistent and noted that models trained with massive text datasets are likely to have extremist and biased content fed into them. In order to deal with bias found in large language models, you can do a post-factor filtering approach like OpenAI does today, but he said in his experience that leads to innocuous things that have nothing to do with Muslims getting flagged as bias, which is another problem
Did nobody think about the consequences of training AI on humanity’s worst impulses? If you input abusive and harmful datasets, you will get abusive and harmful outcomes.
Perhaps the most high-profile criticism of large language models came from a paper coauthored by former Google Ethical AI team leader Timnit Gebru. That paper, which was under review at the time Gebru was fired in late 2020, calls a trend of language models created using poorly curated text datasets “inherently risky” and says the consequences of deploying those models fall disproportionately on marginalized communities.
So somebody DID foresee this problem and they were ignored…