Articles > Information Technology > How do large language models work?
Written by Beth Earnest
Reviewed by Kathryn Uhles, MIS, MSP, Dean, College of Business and IT
If one time-travels to the 1800s in a car, the people they meet would view their vehicle as magical. Of course, it would be merely an example of advanced technology.
Similarly, artificial intelligence is still viewed by some as mysterious and miraculous. Yet it, too, is simply a case of how far the latest technology has come.
One of the fastest-evolving forms of AI is the large language model (LLM), a category of deep learning that uses statistical prediction to generate content. While it has only become widely used within the past five to seven years, this type of system actually emerged decades ago.
In 1966, Massachusetts Institute of Technology computer scientist Joseph Weizenbaum developed , the first program that used natural language programming. It employed “active listening†to reflect a person’s words back to them so it could carry the conversation forward.
In the 1980s, IBM pioneered small language models that could predict the next word in a sentence.
The World Wide Web became available to the public in the early 1990s, and it created an opportunity for language models to access large amounts of data. Deep learning, a form of machine learning, became more prominent in the 1990s, and it was revisited and more heavily implemented in 2011.
By 2022, OpenAI released ChatGPTâ„¢, enabling consumers to converse with a chatbot that used normal, humanlike language. Large language models had suddenly become the norm, not the exception.
While some people might use “large language model†and “generative AI†interchangeably, they are not exactly the same thing. An LLM is a type of generative AI that is focused on processing, understanding and generating human language. All LLMs are generative AI, but not all generative AI are LLMs.
Generative AI:
Large language models:
So, which of these types is ChatGPT? The answer is both. This generative AI chatbot is powered by an LLM, which means it’s trained on text so it can understand context and create humanlike language. It can also convert nontext inputs (like a drawing, for example) into, say, an essay about what kind of art that drawing is.
These models work by mining billions of datasets, breaking them down into smaller pieces, and then feeding the “tokens†into the network so it can learn from them.
As more people begin to use AI as part of their daily routine, both individuals and businesses have found unique uses for the models.
Those uses include:
Language models are trained on a massive language dataset — much of which comes from the internet. Some models use some or all of three stages: pretraining, custom data training and fine-tuning.
In this stage, developers feed the model a large amount of texts from books, articles and websites. This gives the model the capability to process and generate language in a wide range of circumstances.
However, before the model can begin to learn from the data, developers need to “remove the noise,†deleting nontext elements of the information and duplicates. They often use AI to automatically accomplish this goal.
As previously noted regarding translation, after removing the noise, the models convert the data into a numerical format, called tokenization. They then feed the datasets into the network so they can learn how to predict which word comes next in a sequence.Â
Pretraining works well as a foundation, but when developers want a large language model to perform a specific task for an organization, they need to train it with custom data. One way of doing this is by working from scratch with custom data.
One prominent example of a customized model is Med-PaLM 2Tâ„¢. Google developed this model from research and clinical data to answer complex medical questions.Â
This is the process of taking a pre-trained model and training it on a specific dataset. Essentially, it’s taking the best aspects of both of the pretraining and custom data training stages. It can allow the model to go beyond general usage and align with an organization’s goals and objectives.
People seeking careers in the technology sector may want to consider learning how to both use and program these models. Several types of roles require such skills:
For anyone who wants to explore large language models and how they can be used in different industries, °®ÎÛ´«Ã½ offers several information technology programs, including:
For more information about these and other technology degree programs, contact °®ÎÛ´«Ã½.
ChatGPT is a trademark of OpenAI OpCo, LLC.
Med-PaLM 2 is a trademark of Google LLC.
LangChain is a registered trademark of LangChain Inc.
A former newspaper journalist, Beth Earnest has more than 25 years of experience as a professional writer. She has worked with healthcare systems, insurance companies, nonprofits and educational institutions.Â
Currently Dean of the College of Business and Information Technology, Kathryn Uhles has served °®ÎÛ´«Ã½ in a variety of roles since 2006. Prior to joining °®ÎÛ´«Ã½, Kathryn taught fifth grade to underprivileged youth in °®ÎÛ´«Ã½.
This article has been vetted by °®ÎÛ´«Ã½'s editorial advisory committee.Â
Read more about our editorial process.
Learn how 100% of our IT degree and certificate programs align with career-relevant skills.
Download your pdf guide now. Or access the link in our email.