Skip to main content
News Directory 3
  • Business
  • Entertainment
  • Health
  • News
  • Sports
  • Tech
  • World
Menu
  • Business
  • Entertainment
  • Health
  • News
  • Sports
  • Tech
  • World
Revolutionizing LLMs: New Compression Technique Boosts Efficiency and Privacy - News Directory 3

Revolutionizing LLMs: New Compression Technique Boosts Efficiency and Privacy

November 18, 2024 Catherine Williams Tech
News Context
At a glance
Original source: engineering.princeton.edu

Large language models (LLMs) automate tasks such as translation and customer service. However, using LLMs often involves sending requests to centralized servers, which is costly and slow.

Researchers have now introduced a new technique to compress LLMs, which can enhance privacy, save energy, and reduce costs. This algorithm, developed by Princeton and Stanford engineers, trims redundancy and lowers precision within the LLM’s data. A compressed LLM can be stored and accessed locally on devices like smartphones and laptops while maintaining accuracy.

Study coauthor Andrea Goldsmith states that reducing computational and storage demands enables AI on devices that couldn’t otherwise manage such tasks. Coauthor Rajarshi Saha emphasizes the cost of sending requests to backend servers and advocates for LLM inference using consumer GPUs, made possible through compression.

Their algorithm, CALDERA (Calibration Aware Low precision DEcomposition with low Rank Adaptation), will be presented at the upcoming NeurIPS conference. The researchers began their work on compressing large data sets before applying it to LLMs.

LLMs consist of weight matrices, which represent learned word patterns. Saha notes that they developed a generic compression algorithm that could apply to both data sets and models. Their approach combines low-precision and low-rank techniques for greater efficiency than either method alone.

What are the key benefits of compressing large language models for consumer devices?

Interview with Dr. Andrea Goldsmith on the Breakthrough in Large Language Model Compression

NewsDirectory3: Good afternoon, Dr. Goldsmith. Thank you for speaking with us today about your innovative work on compressing large language models. Can you start by explaining the main challenge that current LLMs face when it comes to deployment on consumer devices?

Andrea Goldsmith: Good afternoon, and thank you for having me. The primary challenge with current large language models—LLMs—is their reliance on centralized servers for processing. This leads to high costs, significant latency, and potential privacy concerns. Users often must send their data over the internet to utilize these models, which can be slow and may expose sensitive information.

NewsDirectory3: Your team has developed a new compression algorithm. Can you elaborate on how this algorithm works and the specific improvements it brings?

Andrea Goldsmith: Certainly! Our algorithm is designed to enhance the efficiency of LLMs by trimming redundancy and lowering precision in the data they use. This results in a smaller model that requires less computational and storage power to operate. As a result, a compressed LLM can be stored directly on devices like smartphones and laptops. Despite these reductions, we’ve worked hard to ensure that the model’s accuracy remains intact, allowing users to benefit from AI capabilities without the trade-offs traditionally expected.

NewsDirectory3: What are the implications of being able to run LLMs locally on devices?

Andrea Goldsmith: Running LLMs locally opens up numerous possibilities. First, it enhances user privacy since data doesn’t need to leave the device. Second, it can significantly reduce energy consumption because users won’t have to rely on energy-intensive cloud processing. Additionally, it can lower operational costs for businesses by minimizing reliance on server resources. This democratizes access to AI technologies, enabling even low-resource devices to leverage sophisticated AI tools.

NewsDirectory3: How do you see this technology impacting sectors such as customer service and translation, which rely heavily on LLMs?

Andrea Goldsmith: The impact could be transformative. In customer service, companies can use compressed LLMs to power chatbots and virtual assistants directly on devices, enhancing response times and improving customer satisfaction. In the translation space, users can access real-time translation tools without needing constant internet connectivity, making communication more seamless across the globe. This technology can provide faster, more efficient, and ultimately, a more user-friendly experience.

NewsDirectory3: Are there any potential drawbacks or limitations to using compressed LLMs?

Andrea Goldsmith: Every technology has its limitations. While our approach significantly reduces the size and computational needs of LLMs, it may not be suitable for every application—especially those requiring the very highest levels of precision. Furthermore, there is always a need for continuous monitoring and evaluation to ensure the models remain unbiased and accurate as they evolve with localized use. However, our approach certainly represents a significant step forward.

NewsDirectory3: As this technology begins to roll out, what do you envision for the future of AI in mobile devices?

Andrea Goldsmith: I envision a future where AI is an integral part of everyday mobile experiences. Imagine having powerful language understanding, personalized recommendations, and assistance tools readily available on your phone, with almost no latency and minimal resource usage. This could more deeply integrate AI technologies into our lives, enhancing everything from productivity to creativity while maintaining user control over their data.

NewsDirectory3: Thank you, Dr. Goldsmith, for sharing these exciting developments with us today. We’re looking forward to seeing how your work continues to evolve!

Andrea Goldsmith: Thank you for having me. I’m excited about the future and the role our research can play in making AI more accessible and efficient for everyone.

The team tested CALDERA with open-source models Llama 2 and Llama 3, achieving a performance improvement of up to 5% in predicting word sequences. They evaluated the models using various benchmark tasks, such as logical ordering and physical reasoning questions.

Goldsmith expressed satisfaction with the results, highlighting that focusing on weight matrices improved performance. Compressed LLMs are suitable for tasks that don’t need high precision. This approach also improves privacy by allowing model adjustments on personal devices without sharing data with third parties, reducing risks like data breaches.

However, Saha warned that running LLMs on smartphones or laptops could drain battery life. Low-precision computation helps with power consumption but is not a catchall solution. The proposed techniques can work together to make LLM use on mobile devices more efficient.

The research paper detailing this work will be presented at NeurIPS in December 2024. Along with Goldsmith, Saha, and Mert Pilanci, coauthors include researchers Naomi Sagan and Varun Srivastava. Support came from the U.S. National Science Foundation and other organizations.

Share this:

  • Share on Facebook (Opens in new window) Facebook
  • Share on X (Opens in new window) X

Keep reading

  • DHS Used Palantir System to Profile First Amendment Observers
  • Berck and Calais Face Off in France’s Favorite Beach Final

Related

Search:

News Directory 3

News Directory 3 catalogs US newspapers, news services, newsstands and digital news outlets across all 50 states. Browse local publishers by city, state, or topic, and follow current headlines linked back to their original sources.

Quick Links

  • Disclaimer
  • Terms and Conditions
  • About Us
  • Advertising Policy
  • Contact Us
  • Cookie Policy
  • Editorial Guidelines
  • Privacy Policy

Browse by State

  • Alabama
  • Alaska
  • Arizona
  • Arkansas
  • California
  • Colorado

© 2026 News Directory 3. All rights reserved.
For contact, advertising, copyright, issues email: office@newsdirectory3.com