November 11, 2024

Opinion: The Copyright and Ethical Situation of AI and Journalism

By Taeya Borek
@redlineproject

Art direction by Taeya Borek | Illustration by MidJourney

AI Disclosure: ChatGPT was used to outline a portion of this paper and to generate research on the topic. MidJourney was used to create images. The summary was generated by Summarize Wise custom GPT. The podcast and FAQ were generated by NotebookLM. The video was generated by Lumen5.

Summary: The paper explores the impact of artificial intelligence (AI) on journalism, focusing on copyright and ethical issues. It argues that AI models should not use writers’ works for training without permission or compensation, as this usage does not qualify for “Fair Use” due to its commercial nature and potential harm to the market for original content. Ethically, AI’s bias, inaccuracies, and potential to produce misleading information — such as deepfakes — pose risks to journalistic integrity. The author suggests that AI’s role in journalism be limited to avoid undermining ethical standards and advocates for stronger protections for writers’ intellectual property.

Artificial Intelligence, or AI, used to be known to most through movies and television, usually portrayed as an evil, human-made creation on a mission to destroy the human race. While the latter has yet to come true, AI is no longer an element of fiction, it is real, it is here, and it is revolutionizing industries across the globe. Journalism is one field that has been most impacted by the quick development of AI. It has allowed journalists quick access to summarization, editing, AP style checksheadline creation, and research on specific topics. However, it is not without flaws.

The use of AI, as well as the software itself, has raised copyright and ethical concerns in the writing world. Among these are the fact that AI models are trained on copyrighted works and that AI lacks standards that contribute to ethical journalism. This paper will further explore these concerns, and argue that AI tools — and specifically large language models (LLMs) — should not be allowed to use writer’s work to train their chatbots without their permission or compensation for it. It will be shown that this use does not fall under the Fair Use section of Copyright Law due to commercial use, excessive use of copyrighted material, and impacts on the market. Furthermore, this paper will demonstrate that AI should not be used in the writing of publications because of its bias and discrimination, lack of accuracy, and production of deep fakes.

Copyright Conundrum

As AI models continue to evolve and grow, so too does the amount of copyrighted material it uses to train its chatbots. However, unless AI companies start compensating writers, or gaining their permission, they should not be able to use their material. This is because this use is not protected by the Fair Use exception of Copyright Law.

Copyright Law protects authorship of original works varying from written material to designs, photos, movies, and more. Under this law, creators can publish their material without worry that their work will be stolen or reproduced. Unfortunately, because of AI that is exactly what is happening. AI chatbots are using extensive amounts of copyrighted material, including articles, books, etc, to train their systems. From these materials, the models learn how to produce natural human language, and the creators of AI systems proclaim a section of Copyright Law allows them to do so.

That is section 107 of the Copyright Act, the Fair Use doctrine. The Fair Use section gives allowance for the use of copyrighted works for certain purposes. According to the U.S. Copyright Office, these purposes include criticism, comment, news reporting, teaching, scholarship, and research. Additionally, the office states when analyzing what counts as fair use, courts look to important factors, like purpose and character of the use and amount of the copyrighted work used.

Currently, lawsuits are underway on this topic. One huge lawsuit will most likely set the standard for copyright law and AI going forward, The New York Times v OpenAI and Microsoft. According to a New York Times article from last year, they are arguing that millions of their articles have been used to train OpenAI’s chatbots, like ChatGPT, that now compete with their site for reliable information. However, until a decision is reached, we must rely on the facts that show AI’s use of copyrighted material does not fall under the Fair Use doctrine.

PODCAST: Listen to an AI-generated podcast about the copyright and ethical issues covered in this paper.

One reason to support this is commercial use. As mentioned previously, courts look to specific factors when determining whether fair use can be validated. One of those factors is purpose and character, or determining how a party is using the copyrighted work, which includes nonprofit educational purposes or commercial. Courts tend to favor non-commercial use, and are less likely to grant fair use for commercial purposes. However, AI is using copyrighted material to make money. According to a 2023 article by Copyright Alliance, popular AI platforms like Midjourney, Dal-e, and ChatGPT, offer commercial subscription plans, with more platforms planning to soon. While it is not stated how much of their revenue is generated from its subscriptions, Business of Apps reported that OpenAI, ChatGPT’s owner, generated $1.6 billion in revenue in 2023. These sites are making commercial profit off copyrighted material that they do not own, and are not compensating those who have created the work. In reality, this use is commercial exploitation, as AI sites benefit immensely from writer’s work, while the writers themselves never see a dime.

Another factor courts consider is the amount of copyrighted work a party has used. By the Copyright Office’s standards, fair use is less likely to be found if a group uses a large amount of copyrighted work, especially if they use the entirety of that work, which is what AI sites do. To train their models, “AI generators must copy as much as possible from expressive works, including the most expressive or crucially creative parts of the work, to achieve the purpose of training to generate quality output.” as stated by the Copyright Alliance article. Since AI programs are copying much more than the normal amount granted for fair use, they do not fall under this doctrine of exceptions.

The most important aspect that must be considered is AI’s impact on the market and value of the copyrighted work it uses. This is an additional factor courts examine for fair use, and in relation to this topic, is the most consequential. If there is a high level of, or even any, findings that the unlicensed use of material may harm the future or existing market in which the copyrighted work belongs to, there is a slim to none chance a party will be granted fair use. This is because the use becomes a market substitute for the copyrighted work, and this is the primary reason AI sites do not qualify for fair use.

Their ability to produce everything from news articles to blog posts to books, makes them a competitor in the market of which the copyrighted material used to train their chatbots comes from. This could lead to AI programs generating more revenue than, for example, news outlets like The New York Times, while simultaneously not compensating them for their work being used in their systems. These reasons demonstrate that, by the U.S. Copyright Office’s factors, the use of copyrighted material to train AI models does not qualify for Fair Use, and validates the statement that AI sites must gain the writer’s permission, or give compensation, in order to use a writer’s work in the training of their models.

The Ethics of Artificial Intelligence

In addition to copyright concerns, the journalism field is also facing an ethical dilemma with the use of AI. AI programs offer a vast array of tools that would appeal to any writer. They can generate prompts, offer sources for research, fix mistakes in spelling, and so much more. While this use is okay, so long as the writer identifies what AI was used for, it should never be used for the writing aspect of published work. This is due to the simple truth that AI is inherently unethical, and its use in writing may undermine the journalistic value of a piece.

First and foremost, there are certain standards a journalist must follow when reporting and writing. One of the most important being to remain unbiased and to never discriminate. However, AI has proven numerous times and in numerous different apps that it not only has biases, but it is sexist and racist. According to a 2016 New York Times article, in the year prior Google’s photo app, that uses AI, automatically applied labels to pictures and “was classifying images of black people as gorilla.s” In 2018, Amazon had to get rid of its AI-powered recruiting tool that was designed to rank the best candidates that applied. Why? According to BBC News, it was because it was biased toward women and would penalize resumes, or lower rank them, if they contained words such as ‘women’ or ‘female’. These are only a few examples of the many that occur because AI was programmed primarily by white men on data sets primarily written by white men. These biases are being built into the machines, and therefore AI can not be trusted, as it stands now, to write unbiased, objective material.

Art direction by Taeya Borek | Illustration by MidJourney

AI should also not be used in published writing due to its accuracy problem, or its production of hallucinations. A hallucination in AI is when chatbots produce information that is either entirely false, or misleading. When this occurs, the information may look real, as some of it may be, but when fact-checked turns out to be partly or entirely incorrect. It can happen with any type of information, from dates of events to medical explanations to quotes. According to MIT Management, AI was used in a real-life court case, Mata v Avianca, by a New York lawyer to conduct his legal research. The judge found that citations and quotes used by the attorney were not real, ChatGPT had completely made them up. It does this because AI does not analyze what is factual and what is not, it focuses on producing human language despite if it is truthful. In journalism being accurate comes above all else, and any errors can produce distrust in the media. If writers were to use AI generated content in their work, or for the whole if it, inaccurate information could be spread, discrediting authors or entire media outlets.

On the topic of false information, another consequence from using AI in written works is deepfakes. As in most journalistic pieces, videos, pictures or audio are used to convey the story better. AI now has the power to create its own audio or visuals that sound or look absolutely real, even though they aren’t. These are deepfakes. If journalists were to use these unchecked, unreal sounds, videos or pictures in their pieces they would also be spreading false information through them. A writer could generate a picture of anyone they wanted, say a politician doing something illegal or wrong, then write a story about it (or have AI write it for them), and publish it. In fact, something similar happened. According to a Business Insider article from 2023, a political party from Belgium uploaded a video of Donald Trump giving a speech in 2018. In it, Trump called for Belgium’s withdrawal from the Paris Climate Agreement. The speech never happened though, it was a deepfake. If journalists are to remain authentic in their work, they can not use AI-generated visuals in their publications, as it can perpetuate fake information, be it purposefully or not.

Artificial Intelligence isn’t going anywhere, instead what we’ve seen of it is just the start. However, to ensure journalists and writers don’t go anywhere either we must agree that their work, their copyrighted material, is protected and ensure it is not being used unfairly. Copyright Law was made to protect people and their creations, not technology’s theft of it. Furthermore, if journalists are to remain ethical, accurate, and the voice of truth, as they have been for generations, they need to refrain from using AI to write their work or create visuals for it. Going forward, AI companies should seek permission from authors or compensate for their material in the training of AI models, and journalists must limit the role AI plays in their content. By doing so, we will ensure that writers continue to do what they do best; write.

AI in Journalism: Copyright and Ethical Concerns FAQ

1. How is AI being used in journalism?

AI is revolutionizing journalism by providing tools for:

  • Summarization: Condensing large amounts of information into concise summaries.
  • Editing: Identifying and correcting errors in grammar, style, and factual accuracy.
  • AP Style Checks: Ensuring adherence to journalistic style guidelines.
  • Headline Creation: Generating catchy and informative headlines.
  • Research: Quickly gathering information on specific topics.

2. What are the main copyright concerns regarding AI in journalism?

The primary concern is that AI models are trained on massive datasets of copyrighted material, including articles and books, without permission or compensation for writers. This raises questions about fair use and the potential for AI to profit from the work of human authors without proper attribution or remuneration.

3. Does the “Fair Use” doctrine apply to AI’s use of copyrighted material?

The paper argues that AI’s use of copyrighted material does not fall under Fair Use for several reasons:

  • Commercial Use: AI companies are profiting from the use of copyrighted material.
  • Excessive Use: AI models copy vast amounts of copyrighted work, often exceeding the limits of fair use.
  • Market Impact: AI-generated content can compete with the original copyrighted works, potentially harming the market for human-authored content.

4. What are the ethical concerns surrounding AI in journalism?

Key ethical concerns include:

  • Bias and Discrimination: AI models can perpetuate existing biases present in the data they are trained on, leading to potentially discriminatory content.
  • Lack of Accuracy: AI models can generate “hallucinations” — fabricated or misleading information — which can undermine the credibility of journalistic work.
  • Deepfakes: AI can create realistic but fake audio and visual content, raising concerns about the spread of misinformation.

5. How can AI bias affect journalistic content?

AI models trained on biased data can produce content that reflects and amplifies those biases. For example, an AI model trained on a dataset with predominantly male authors might produce content that underrepresents female perspectives.

6. What is an AI “hallucination” and why is it problematic for journalism?

An AI hallucination is when a chatbot or language model generates false or misleading information. This is problematic for journalism because accuracy is paramount, and the spread of false information can erode public trust.

7. How can deepfakes impact journalism?

Deepfakes can create fabricated audio and visual content that appears authentic, making it difficult to distinguish real events from fabricated ones. This can lead to the spread of misinformation and damage the credibility of news sources.

8. What are some recommendations for addressing the ethical and copyright concerns related to AI in journalism?

  • Transparency: Journalists should disclose their use of AI tools.
  • Human Oversight: AI should be used as a tool to assist journalists, not replace them.
  • Copyright Reform: Legal frameworks may need to be updated to address the unique challenges posed by AI-generated content.
  • Ethical Guidelines: The journalism industry should develop ethical guidelines for responsible AI use.

—–

Editor’s note: Students in Mike Reilley’s AI Journalism (COMM 294) Fall 2024 undergraduate course experimented with AI storytelling tools to create these stories, following the AI use guidelines on The Red Line Project’s Principles page.

Bookmark the permalink.

Leave a Reply

Your email address will not be published. Required fields are marked *