°®ÎÛ´«Ã½

Skip to Main Content Skip to bottom Skip to Chat, Email, Text

Articles > Information Technology > How do large language models work?

How do large language models work?

If one time-travels to the 1800s in a car, the people they meet would view their vehicle as magical. Of course, it would be merely an example of advanced technology.

Similarly, artificial intelligence is still viewed by some as mysterious and miraculous. Yet it, too, is simply a case of how far the latest technology has come.

How long have large language models been around? 

One of the fastest-evolving forms of AI is the large language model (LLM), a category of deep learning that uses statistical prediction to generate content. While it has only become widely used within the past five to seven years, this type of system actually emerged decades ago.

In 1966, Massachusetts Institute of Technology computer scientist Joseph Weizenbaum developed , the first program that used natural language programming. It employed “active listening†to reflect a person’s words back to them so it could carry the conversation forward.

In the 1980s, IBM pioneered small language models that could predict the next word in a sentence.

The World Wide Web became available to the public in the early 1990s, and it created an opportunity for language models to access large amounts of data. Deep learning, a form of machine learning, became more prominent in the 1990s, and it was revisited and more heavily implemented in 2011.

By 2022, OpenAI released ChatGPTâ„¢, enabling consumers to converse with a chatbot that used normal, humanlike language. Large language models had suddenly become the norm, not the exception.

What’s the difference between an LLM and generative AI? 

While some people might use “large language model†and “generative AI†interchangeably, they are not exactly the same thing. An LLM is a type of generative AI that is focused on processing, understanding and generating human language. All LLMs are generative AI, but not all generative AI are LLMs.

Generative AI:

  • Is a wide category of systems that produce new data
  • Uses not only text but also images, audio, video and computer code as resources
  • Is trained on many kinds of datasets based on the output desired (for example, if a person wants a specific type of art, it will have access to millions of art pieces)

Large language models:

  • Are a subset of generative AI tailored specifically for language
  • Usually use text as resources
  • Are trained mostly on texts, including books, newspapers and websites
  • Create only text outputs; however, they can now accept audio and video inputs and convert them into text outputs

So, which of these types is ChatGPT? The answer is both. This generative AI chatbot is powered by an LLM, which means it’s trained on text so it can understand context and create humanlike language. It can also convert nontext inputs (like a drawing, for example) into, say, an essay about what kind of art that drawing is.

These models work by mining billions of datasets, breaking them down into smaller pieces, and then feeding the “tokens†into the network so it can learn from them.

How are large language models used?

As more people begin to use AI as part of their daily routine, both individuals and businesses have found unique uses for the models.

Those uses include:

  • Answering questions: One of the biggest emerging uses is chatbots on company websites. They use natural language processing (NLP) to analyze text and determine the user’s intent with the question. These, too, have challenges — chatbots have been known to create “hallucinations,†or inaccurate information that they present as known facts. One popular meme illustrates this: A person hunting for mushrooms asks a chatbot if a certain mushroom is safe. After the chatbot confirms it is safe, he eats it. The next frame shows a tombstone, with the chatbot saying, “You’re right — that mushroom was poisonous! I’m sorry for the confusion. Would you like to learn more about poisonous mushrooms?â€
  • Coding: Large language models can write and debug software code across multiple programming languages.
  • Analyzing data: The models can perform analysis, extract specific data points and classify specific sets of data. They can do so faster than a human can, which makes them a valuable asset. Businesses should make sure, however, that a human double-checks the model’s work, looking for possible bias or data that is overlooked.
  • Generating content: With just a small amount of input, models can create articles, scripts, emails, reports and marketing copy. However, as organizations are increasingly learning, when they use AI, they should have a human constantly checking for accuracy. In 2023, newspaper chain suspended its use of an AI service after it made several highly noticeable mistakes. The mistakes led to viral memes that mocked the newspaper company for its carelessness.
  • Translating: Models analyze billions of multilingual texts to provide translations for dozens of languages. They do this by:
    • Breaking down the source text into smaller chunks called tokens and converting those chunks into numbers.
    • Mapping the numerical tokens in a high-dimensional space, where words with similar meanings cluster near one another. This allows the model to figure out the contextual meaning of the phrase.
    • Processing whole sentences at one time using neural networks called transformers.
    • Using the above process to predict the probable next token.

What data do the models use to produce responses to queries?

Language models are trained on a massive language dataset — much of which comes from the internet. Some models use some or all of three stages: pretraining, custom data training and fine-tuning.

Pretraining

In this stage, developers feed the model a large amount of texts from books, articles and websites. This gives the model the capability to process and generate language in a wide range of circumstances.

However, before the model can begin to learn from the data, developers need to “remove the noise,†deleting nontext elements of the information and duplicates. They often use AI to automatically accomplish this goal.

As previously noted regarding translation, after removing the noise, the models convert the data into a numerical format, called tokenization. They then feed the datasets into the network so they can learn how to predict which word comes next in a sequence. 

Custom data training

Pretraining works well as a foundation, but when developers want a large language model to perform a specific task for an organization, they need to train it with custom data. One way of doing this is by working from scratch with custom data.

One prominent example of a customized model is Med-PaLM 2T™. Google developed this model from research and clinical data to answer complex medical questions. 

Fine-tuning

This is the process of taking a pre-trained model and training it on a specific dataset. Essentially, it’s taking the best aspects of both of the pretraining and custom data training stages. It can allow the model to go beyond general usage and align with an organization’s goals and objectives.

The future of large language models 

People seeking careers in the technology sector may want to consider learning how to both use and program these models. Several types of roles require such skills:

  • Software developers: Some developers are expected to build integrations with model application programming interfaces. They also may need to use tools such as ® to build retrieval-augmented generation pipelines.
  • Data scientists: Companies want to fine-tune existing models to fit their needs, and data scientists can help with that.
  • Nontechnical tech roles or product managers: People in these types of jobs may not need to actually develop AI platforms themselves, but they may be called upon to translate AI capabilities into actual value for their business.

Learn more about large language models

For anyone who wants to explore large language models and how they can be used in different industries, °®ÎÛ´«Ã½ offers several information technology programs, including:

For more information about these and other technology degree programs, contact °®ÎÛ´«Ã½.

ChatGPT is a trademark of OpenAI OpCo, LLC.

Med-PaLM 2 is a trademark of Google LLC.

LangChain is a registered trademark of LangChain Inc.

Headshot of Beth Earnest

ABOUT THE AUTHOR

A former newspaper journalist, Beth Earnest has more than 25 years of experience as a professional writer. She has worked with healthcare systems, insurance companies, nonprofits and educational institutions. 

Headshot of Kathryn Uhles

ABOUT THE REVIEWER

Currently Dean of the College of Business and Information Technology, Kathryn Uhles has served °®ÎÛ´«Ã½ in a variety of roles since 2006. Prior to joining °®ÎÛ´«Ã½, Kathryn taught fifth grade to underprivileged youth in °®ÎÛ´«Ã½.

checkmark

This article has been vetted by °®ÎÛ´«Ã½'s editorial advisory committee. 
Read more about our editorial process.

FREE IT Programs Guide

Learn how 100% of our IT degree and certificate programs align with career-relevant skills.

Thank you

Download your pdf guide now. Or access the link in our email.

FREE IT programs guide. Please enter your first and last name.