# Continuum Labs - Applied AI

"Applied Artificial Intelligence"

### <mark style="color:purple;">Our platform</mark>

We have created a platform that makes the process of creating a proprietary neural language model and integrating it into enterprises of all sizes <mark style="color:yellow;">simple and intuitive.</mark>

We have shortened the development time frame down to months.   We can work with you to determine the optimal use case for your AI model and then translate that into a working application rapidly.

From data creation and curation, fine tuning, optimisation and then deployment - the Continuum process is systematic at the same time as being flexible.

We work with all organisation types - from start ups that want to make AI the core of their organisation - to large enterprises and government that want to begin their AI journey.

### <mark style="color:purple;">**Unlocking the Power of Customised Models**</mark>

Our mission is to deliver tailor-made AI applications that cater to your specific domain and use cases.&#x20;

We are committed to democratising access to AI, ensuring that anyone can use the power of this technology.   &#x20;

You do not have to outsource your AI journey to external service providers; partner with us to embark on your unique AI journey.

Our core believe is *<mark style="color:yellow;">**you should own your models**</mark>*, hosted securely on your  own infrastructure.&#x20;

We recommend taking a portfolio approach toward this technology - pick several low cost, low risk use cases - and start today.

### <mark style="color:purple;">Why should I make the investment?</mark>

#### <mark style="color:green;">**Customisation**</mark>

With an in-house model, your organisation gains an understanding of how AI can be used to make the team more competitive, more productive and ultimately please your customers. &#x20;

With an internal model, a company can tailor the model to its specific domain, industry, and use cases.

This allows for fine-tuning the model to better understand and generate content relevant to the company's products, services, and customers.  By continuously adapting the model to the company's evolving needs, it can stay ahead of competitors that are relying on generic, off-the-shelf models.

Off-the-shelf solutions may fall short when it comes to meeting your unique needs. By developing your own AI applications in-house, you gain the flexibility to *<mark style="color:yellow;">**customise it precisely to your specifications,**</mark>* ensuring optimal performance and relevance.

This factors should all combine to generate a competitive edge.

#### <mark style="color:green;">**Data Ownership**</mark>

**This issue cannot be understated.** &#x20;

If you are using external models, *<mark style="color:yellow;">**your data is leaking out of your organisation**</mark>*.  &#x20;

Your organisation's knowledge, internal processes and culture are your core assets - you cannot have security over these assets if your AI infrastructure relies on external models.

With a proprietary model, you retain full ownership and control over your data, safeguarding it against competitive threats. privacy breaches and ensuring compliance with regulatory requirements.

Sensitive information can be kept within the company's infrastructure, reducing the risk of data breaches or unauthorised access.&#x20;

#### <mark style="color:green;">Proprietary knowledge and trade secrets</mark>

The privacy and security an internal model delivers means it can be *<mark style="color:yellow;">**trained on a company's proprietary data,**</mark>* including internal documents, customer interactions, and industry-specific knowledge.&#x20;

This enables the model to capture and leverage the company's unique expertise and trade secrets, providing a competitive advantage that cannot be easily replicated by others.

#### <mark style="color:green;">**Future Development**</mark>

Having full control over your AI models allows *<mark style="color:yellow;">**rapid experimentation**</mark>* with new ideas, architectures, and training techniques.&#x20;

This agility enables faster innovation cycles, allowing the company to explore novel applications and use cases before competitors can catch up.   The ability to iterate and improve the model rapidly can lead to breakthrough innovations and first-mover advantages. &#x20;

#### <mark style="color:green;">**Talent Retention**</mark>

Developing an internal model can attract top talent in the field of deep learning and artificial intelligence.   The opportunity to work on the latest technology and solve unique challenges can be a strong draw for skilled researchers and engineers.&#x20;

By building a talented team around the internal model, a company can foster a culture of innovation and expertise that contributes to an AI driven culture.

#### <mark style="color:green;">**Long-Term Total Cost of Ownership**</mark>

While there is an upfront investment for in-house model development and ownership, it often results in long-term cost savings compared to ongoing licensing fees associated with large, generic off-the-shelf solutions.

By eliminating the need for third-party licenses and subscriptions, you can reduce ongoing expenses. &#x20;

Additionally, an internal model can be optimised for the company's specific hardware and infrastructure, ensuring efficient resource utilisation and minimising waste.

Furthermore, there are many use cases that do not require large, expensive models.  You can use open source models that are small, but when fine tuned and optimised can perform highly valuable tasks withing your organisation. &#x20;

#### <mark style="color:green;">Flexibility and control over ethical considerations</mark>

Owning an internal model allows full control over the ethical considerations and biases embedded within the model.&#x20;

You can ensure that the model aligns with your organisation's values, principles, and ethical guidelines. This level of control is important for maintaining trust with customers, partners, and stakeholders, especially in industries where fairness, transparency, and accountability are paramount.

#### <mark style="color:green;">**Infrastructure Integration**</mark>

An internal model can be seamlessly integrated with a company's existing software systems, workflows, and data pipelines.

This enables the model to leverage real-time data and provide actionable insights directly within the company's applications.  By having control over the model's deployment and integration, a company can create a more streamlined and efficient workflow.

### <mark style="color:purple;">Our offer</mark>

Working with Continuum to build your proprietary application is a forward-thinking investment in the future.  We work with you to become AI ready and your team AI literate.

The end result is a bespoke model tailored specifically to your needs.

You retain control over the data it's trained on, the knowledge base it's integrated with, and the functions it performs.&#x20;

Your data remains secure, shielded from external service providers, ensuring confidentiality and privacy. &#x20;

More importantly, your proprietary model can be continuously refined and enhanced, evolving alongside your evolving needs and use cases.

With ownership comes <mark style="color:yellow;">control over your intellectual property</mark>, granting you the ability to capitalise on your proprietary models through monetisation or licensing opportunities.&#x20;

## <mark style="color:blue;">Quick links</mark>

{% content-ref url="/pages/lkGiHdGvoygkVk231tq3" %}
[What we do](/continuum-applications/overview/what-we-do)
{% endcontent-ref %}

{% content-ref url="/pages/3IoSyxed5OkuHutaXb9E" %}
[Our Features](/continuum-applications/overview/our-features)
{% endcontent-ref %}


# What we do

Continuum AI Driven Modules

### <mark style="color:purple;">Why do we exist?</mark>

We work with organisations to implement and own their own internal large language models.

We create proprietary models that you own and control.   Models that you can customise and integrate with full privacy and security into your workflows.  There is no risk of your data leaving your organisation, there is no risk of your exposing your customers data to external third party models.&#x20;

While generative AI is technology that has the power to transform workflows, lift productivity and deliver competitive edge - it is still a nascent technology.  There are many risks associated with opening your organisation up to third party models that you have no control over and no genuine transparency as to how the work.

The real power of generative AI comes from providing it access to your internal data and workflows.  You cannot enjoy this power fully unless you own and operate.  Apart from the obvious security and competitive issues - as regulation evolves, it will become increasingly difficult to legally allow third party models access to your private data.

These risks do not exist by owning and operating your own large language model.  While there is an investment from you, in time and resources - the total cost of ownership is similar to using third party models.

But the pitch is not just about near term economics - by taking a medium to long term view - you can build internal expertise and valuable intellectual property.  Every organisation should be building towards being AI literate.  This cannot be achieved outsourcing your AI models to third party technology companies.

### <mark style="color:purple;">Productivity benefits</mark>

Using generative AI integrated with your organisation's knowledge can create internal knowledge workers that can increase productivity up to 20%. &#x20;

This is a *<mark style="color:yellow;">**permanent increase in productivity.**</mark>*   Your organisation will be able to generate a major lift in economic output and enhanced decision making - increasing shareholder value.

We believe that the ability to effectively integrate AI into your operations is not in option, but a *<mark style="color:yellow;">**critical source of future competitive advantage.**</mark>* &#x20;

Our core thesis is that organisations that continuously develop and integrate AI models into their operations will generate significant edge within their industry.

<mark style="color:blue;">**What sets us apart is our commitment to making AI accessible and easy to implement.**</mark> &#x20;

You do not need a large budget or an expert internal team to get started.  There is no need to deal with large consultancy firms - it is not that complicated.

Our flexible, modular approach allows you to start small and scale up as your needs evolve.   And our team will be with you every step of the way, providing guidance, support, and training to ensure your success.

Please call us and we can explain the process.

From start-ups to large enterprises, we work with companies of all sizes to identify problems that can be solved - **and opportunities created** - by large language models.

### <mark style="color:blue;">Call 1300 924 656 to organise a free consultation to assess your AI readiness</mark>

***

### <mark style="color:purple;">What is our edge?</mark>

At Continuum, we have spent years developing a skill set that allows us to deliver industry grade proprietary models. &#x20;

Continuum Labs is an organisation that comprises two skill sets:

1. We have <mark style="color:green;">**genuine deep learning and artificial intelligence capability**</mark>.  We are able to build generative AI models and integrate them with knowledge banks to create powerful applications.
2. We have a background in <mark style="color:green;">**industry and capital markets.**</mark>  We understand organisational structure and the sources of competitive advantage.  We know what will move the shareholder value dial.

The combination of these skill sets allows us to create AI applications that solve real problems.  &#x20;

We understand how competitive edge is generated - and can design and implement AI applications that provide your business with an advantage.

***

### <mark style="color:purple;">Team Skill Set</mark>

* PhD mathematicians who bring a deep understanding of the theoretical foundations of AI and can develop novel algorithms tailored to your specific use case.
* Deep learning experts with years of experience in designing, training, and optimising neural networks for maximum performance and efficiency.
* GPU infrastructure experts who can build and maintain the high-performance computing environments necessary for large-scale AI deployment, ensuring your models run smoothly and reliably.
* Integration experts who can integrate your custom AI solutions with your existing software systems, workflows, and data pipelines, minimising disruption and accelerating time-to-value.

What sets us apart is not just the depth of our expertise, but also the breadth of our capabilities.

By bringing together specialists from multiple disciplines, we can provide end-to-end support for your AI initiatives, from idea generation and development to deployment and ongoing optimisation.

We have developed proprietary methodologies and tools that streamline the AI development process, allowing us to deliver results faster and more cost-effectively than our competitors.&#x20;

With Continuum, you're not just getting access to AI technology – you're partnering with a team of experts who who have the skills and experience to turn bring AI into your organisation.

***

### <mark style="color:purple;">Our process</mark>

We view this technology as a means to fundamentally change how decisions are made and how people interact with tools and data. &#x20;

We know the capabilities *<mark style="color:yellow;">**as well as the limitations**</mark>* of AI models - we can help you find a use case that will add value immediately.

Simply going through the education process will provide you hands on experience as to how these models work, what they can do.  Your team will come up with the use cases, your creativity will drive the process.

AI can deliver a new way of decision making, but with that comes cultural change and a shift in workplace practices.  The best way to start is slowly.

The first step is to find an initial use case that is:

### <mark style="color:green;">1. Easy</mark>

### <mark style="color:green;">2. High Value</mark>

### <mark style="color:green;">3. Low risk</mark>

### <mark style="color:green;">4. Low cost</mark>

Our experience *<mark style="color:yellow;">**it is not difficult to find a range of use cases that meat all of these criteria.**</mark>* We find every organisation has a pain point that generative AI can be applied to.

We can have an initial consultation that will demystify artificial intelligence and help you understand *<mark style="color:yellow;">**what it can and cannot do**</mark>*. &#x20;

And we can almost be certain, that at the end of the conversation, we will be able to find an interesting use case you can get started on immediately.

### <mark style="color:purple;">What we are not</mark>

We create <mark style="color:blue;">**proprietary models**</mark> customised for you and your organisation.

We do not take off the shelf models, create a user interface and then present it as a unique and 'proprietary' solution.  &#x20;

When dealing with any AI offering, <mark style="color:green;">**look under the hood**</mark> - if it is an application built on top of an external model - it is built on a shaky foundation.   These are models that are never yours, and cannot be developed into a valuable organisational asset.

If you want to genuinely use AI within your organisation - *<mark style="color:yellow;">**you**</mark>* *<mark style="color:yellow;">**cannot rely on external models**</mark>* - apart from the privacy and security issues - you are missing the opportunity to build your own internal AI capability.

<mark style="color:green;">**Customisation**</mark> is also important.  We do not offer generic 'off the shelf' models'. &#x20;

All organisations have distinct processes that provide them with a form of competitive edge.  Trying to offer a one size fits all AI driven application is not going to integrate and support those organisational processes.

Our models are created for your unique use case.

***

### <mark style="color:purple;">What is agency?</mark>

Your AI models can be created to act as a important member of your organisation - enabled with organisational knowledge, an understanding of your operations and its processes and culture.  &#x20;

But the aspect that delivers the most is their ability for <mark style="color:blue;">**agency.**</mark>

The term "agency" refers to the capacity to act and make decisions in the world.&#x20;

AI demonstrates agency by executing tasks that alter the environment or influence outcomes in its domain of operation.  This is a powerful feature, that must be carefully controlled with the appropriate guardrails.

***

### <mark style="color:purple;">Practical Application</mark>

The practical application of AI within any organisation must begin slowly and be tightly controlled.

Privacy and security is paramount.  It is not advisable to allow members of your organisation to use private models such as ChatGPT.   Your data is not just yours, *<mark style="color:yellow;">**but also your customers**</mark>*.  <mark style="color:green;">**Using external models is fraught with regulatory danger.**</mark>

The other primary reason to create your own AI models is to begin to create internal AI knowledge, and allow future optionality.  If you use third party AI models - you lose this optionality.


# Our Features

Rapid development of customised models for immediate deployment

We understand that investing in a custom AI solution can feel like a daunting prospect, especially if you're new to the technology.&#x20;

You may be concerned about the upfront costs, the complexity of the development process, or the potential risks involved.  At Continuum, we can put those concerns to rest.

### <mark style="color:purple;">The Investment</mark>

Developing a custom AI model requires an initial outlay and investment in your time and focus. - but it's important to consider the long-term benefits.&#x20;

By owning your own AI infrastructure, you can avoid the ongoing licensing fees and dependencies associated with third-party solutions.  You'll have *<mark style="color:yellow;">**complete control over your data and intellectual property,**</mark>* and you'll be able to adapt and scale your AI capabilities as your needs evolve.&#x20;

In short, investing in a custom AI solution is an investment in your organisation's future competitiveness and growth.

### <mark style="color:purple;">Complexity</mark>

As for the complexity of the development process, that's where our expertise comes in.&#x20;

AI can seem like a black box, with complex algorithms and technical jargon that can be meaningless to executives.  That's why we've developed a transparent, collaborative approach that keeps you informed and involved at every step of the way.&#x20;

Our team will work closely with you to understand your requirements and translate them into a clear, actionable development plan. &#x20;

We'll break the process down into manageable phases, with regular checkpoints and opportunities for feedback and refinement.  And we'll provide ongoing support and training to ensure that your team is fully equipped to leverage your new AI capabilities.

### <mark style="color:purple;">Data Security</mark>

All of our solutions are designed with strict data protection protocols in place, and we work closely with our clients to ensure compliance with relevant regulations and standards.&#x20;

We also offer flexible deployment options, including on-premises installation and private cloud hosting, so you can choose the approach that best meets your security and infrastructure needs.

At Continuum, our goal is to make the process of adopting AI as seamless and stress-free as possible.&#x20;

We're here to answer your questions, address your concerns, and provide the guidance and support you need to unlock the full potential of this transformative technology.&#x20;

&#x20;With our expertise and partnership, you can confidently embark on your AI journey and stay ahead of the curve in an increasingly competitive landscape.

Call us on <mark style="color:blue;">**1300 924 656**</mark> for a free consultation.

### <mark style="color:purple;">Rapid Development</mark>

We can train and fine tune neural language models and deploy them rapidly. &#x20;

We have a systematic but flexible process for each component of model development:

1. Data creation and curation
2. Fine Tuning&#x20;
3. Testing and Evaluation
4. Front End Development
5. Deployment

At the end of this process, you will have your own internal model.

### <mark style="color:purple;">Customisation</mark>

Your model will be trained exactly how you want it.   We combine latest fine tuning techniques and model integration with organisational knowledge.

Your model can be customised for:

1. Domain knowledge
2. To execute a specific task
3. To behave in a particular manner
4. To be on brand, to mirror your organisational personality

You will be surprised with what these models are capable of becoming.  They can become a part of your team - a peculiar part of your team with a distinct personality and knowledge base. &#x20;

### <mark style="color:purple;">Low cost deployment</mark>

We can deploy your model on a private and secure NVIDIA GPU cluster in Australia.   Your model is available via dedicated endpoints.

Continuum utilise NVIDIA's TensorRT-LLM framework and the NVIDIA Triton Inference Server to develop rapid inference rates at low cost - without compromising on performance.

Your users will experience lightning-fast model responses and low latency.  And most importantly - inference is low cost - so you do not need many users to make an economic use case.

Importantly, our systems scale to meet your growing demands, ensuring consistent performance. &#x20;

As the value of your models  in the organisation improves, and the use of the model increases - we can scale with you.


# Secure and Private GPU Infrastructure

Our core value is data security and privacy.

That's why we have invested in our own private and secure GPU cluster, designed with robust controls to mitigate the risks associated with shared GPU environments.

### <mark style="color:purple;">Key Mitigation Strategies</mark>

<mark style="color:blue;">Physical Partitioning</mark>

Our GPUs are physically partitioned to prevent any cross-process leakage or unauthorised access between different models or clients. This eliminates the risk of attackers stealing sensitive data like model weights or reconstructing model outputs.

<mark style="color:blue;">No Multi-Tenant Access</mark>

We strictly prohibit multi-tenant access to our GPUs. Each client's models and data are isolated in their own secure partition, inaccessible by any other parties.

<mark style="color:blue;">Proactive Patching</mark>

We work closely with our GPU vendors to ensure we rapidly deploy the latest security patches and firmware updates across our entire cluster. Our dedicated security team continuously monitors for any newly discovered vulnerabilities.

<mark style="color:blue;">Secure Coding Practices</mark>

All of the code we develop to train and deploy models on our GPU cluster adheres to rigorous secure coding standards. Our developers are highly trained in identifying and preventing any potential exploit vectors.

<mark style="color:blue;">Ongoing Security Research</mark>

We actively participate in AI security research to stay on the cutting edge of identifying and mitigating GPU and model vulnerabilities. Our team collaborates with leading experts to develop new security measures.

By maintaining our own private and secure GPU infrastructure, with state-of-the-art controls and active risk mitigation, Continuum Labs ensures that our clients' valuable models and data remain fully protected at all times.&#x20;


# Generative AI Implementation Risks

<table><thead><tr><th width="171">Risk Category</th><th width="218" align="center">Description</th><th width="166" align="center">Potential Impact</th><th>Mitigation with Continuum</th></tr></thead><tbody><tr><td>Data Privacy and Confidentiality</td><td align="center">Inadvertent sharing of confidential or private information with GenAI systems</td><td align="center">Legal liabilities, regulatory penalties, reputational damage</td><td>Continuum's secure model hosting and data handling practices ensure strict control over sensitive information</td></tr><tr><td>Legal and Regulatory Compliance</td><td align="center">Questions around ownership of generated content and potential liabilities</td><td align="center">Legal disputes, financial penalties, reputational harm</td><td>Continuum stays up-to-date on evolving regulations and provides guidance on compliant use of GenAI</td></tr><tr><td>Insecure Code Generation</td><td align="center">Reliance on untested AI-generated code introducing vulnerabilities</td><td align="center">Data breaches, system compromises, operational disruptions</td><td>Continuum's rigorous testing and validation processes ensure the security and reliability of generated code</td></tr><tr><td>Trust and Reputation</td><td align="center">Inaccurate or biased GenAI outputs published under company name</td><td align="center">Loss of customer trust, damage to brand reputation, financial losses</td><td>Continuum's custom model training and output monitoring mitigate the risk of inaccurate or biased results</td></tr><tr><td>Workflow Disruption</td><td align="center">GenAI changing workflows and being used by employees in various roles</td><td align="center">Inconsistent practices, decreased productivity, security gaps</td><td>Continuum works closely with clients to integrate GenAI into workflows while maintaining security and efficiency</td></tr><tr><td>Prompt Injection Attacks</td><td align="center">Malicious prompts manipulating GenAI systems to produce harmful outputs</td><td align="center">Data leakage, system compromise, reputational damage</td><td>Continuum implements robust prompt filtering and validation to prevent prompt injection attacks</td></tr><tr><td>Voice Spoofing Attacks</td><td align="center">Synthetic voice generation used for impersonation and fraud</td><td align="center">Financial losses, reputational harm, erosion of trust</td><td>Continuum develops advanced detection capabilities to identify and prevent voice spoofing attacks</td></tr><tr><td>Model Bias and Fairness</td><td align="center">GenAI models reflecting societal biases or discriminating against certain groups</td><td align="center">Legal liabilities, reputational damage, erosion of public trust</td><td>Continuum employs rigorous testing and auditing to identify and mitigate model biases</td></tr><tr><td>Lack of Interpretability</td><td align="center">Difficulty understanding and explaining GenAI decision-making processes</td><td align="center">Regulatory non-compliance, lack of accountability, erosion of trust</td><td>Continuum prioritizes interpretability and provides clear explanations of model outputs</td></tr><tr><td>Insider Threats</td><td align="center">Malicious insiders exploiting GenAI access for unauthorized purposes</td><td align="center">Data theft, system sabotage, reputational harm</td><td>Continuum implements strict access controls and monitoring to detect and prevent insider threats</td></tr></tbody></table>


# Model Range

Our range of models

## <mark style="color:blue;">The Continuum Range of Models</mark>

### <mark style="color:purple;">The Basics</mark>

Continuum can provide a range of models for rapid deployment into your organisation or integration into your infrastructure or application.

These models are across a range of domains, areas of expertise, use case and application.

All models can be further customised for your specific needs.

Note - these models do not replace expert advice, they are a complement to knowledge discovery and to reduce information asymettry.&#x20;

### <mark style="color:purple;">Creating a Project</mark>

The process is simple.  We spend time with you to learn your specific objectives.&#x20;

We can provide you with a comprehensive overview of how the process works.  From data creation and curation, model training, inference and deployment.

### <mark style="color:purple;">The Continuum Model Range</mark>

<table data-full-width="false"><thead><tr><th width="228">Domain</th><th width="165" align="center">Expertise</th><th align="center">Summary</th></tr></thead><tbody><tr><td>Investment Management</td><td align="center">Fundamental Analysis</td><td align="center">The investment management module can become expert on any industry sector - analysing  in real time company news flow and broker intelligence</td></tr><tr><td>Employment Law</td><td align="center">Employment and Awards</td><td align="center">The employment law assistant can be used by employers and employees to navigate the complexity of Australian employment regulations</td></tr><tr><td>Home Insurance</td><td align="center">Insurance</td><td align="center">The home insurance model allows users to ask questions about their specific situation and requirements and have them answered and suggestions provided to help them</td></tr><tr><td>Government Grants</td><td align="center">Federal and State</td><td align="center">The government grant model rapidly analyses in real time government grants and tenders.  It can then provide a framework for grant application, adhering to the specific requirements of the grant or tender</td></tr><tr><td>Consumer Surveying</td><td align="center">Profiling Consumer Insights</td><td align="center">Profiling consumers is important for all brands, the consumer profiling model is trained to understand your brand and engage in long form conversations with consumers to understand their perception and buying intentions</td></tr><tr><td>Consumer Mortgages</td><td align="center">Finance</td><td align="center">The consumer mortgage model provides users with the ability to ask questions about their situation and goals and provide advice on strategy</td></tr><tr><td>Renewable Energy</td><td align="center">Home Appliances</td><td align="center">The renewable energy model provides users with an interface to learn about the myriad of options to install solar and battery technology into their residential property</td></tr><tr><td>Travel</td><td align="center">Europe</td><td align="center">The travel model is an expert on Europe.  It provides users the mechanism to learn about unique European travel destinations, creative travel plans and ultimately design an itinerary</td></tr><tr><td>Search</td><td align="center">Enterprise Search</td><td align="center">The enterprise search model is one of our most powerful.  These models are heavily customised for each organisation to allow users to find anything they want to about their business.</td></tr><tr><td>Aged Care</td><td align="center">Aged Care Schemes</td><td align="center">The aged care model allows users to navigate the highly complex aged care schemes in Australia</td></tr><tr><td>Psychologist</td><td align="center">Personal</td><td align="center">The psychologist model is a personalised model that allows users to ask questions about life, mental health, wellbeing and living well</td></tr><tr><td>Pharmaceuticals</td><td align="center">Healthcare</td><td align="center">This model is an interface into the highly complex Australian pharmaceutical benefits scheme (PBS)</td></tr></tbody></table>

Each of these models can be customised to meet varying objectives. &#x20;

All models are integrated with neural (vector) databases to ensure up to date knowledge and ensure information retrieval is highly accurate.


# Investment Management

A model to give all investors edge

The Continuum team aims to develop IM-GPT, a large language model application designed to serve as an <mark style="color:green;">**assistant investment manager.**</mark>&#x20;

IM-GPT, or Investment Manager-GPT, represents the team's inaugural endeavour in creating a sophisticated tool tailored for the financial industry.&#x20;

### <mark style="color:purple;">What does it do?</mark>

This application offers users insightful commentary and research guidance across various financial markets, including equities, commodities, currencies, and fixed interest securities.&#x20;

IM-GPT has the capability to <mark style="color:green;">**process streaming news flow**</mark>**,** discern crucial information within it, and subsequently analyse this content to provide users with valuable research ideas and perspectives.&#x20;

IM-GPT is <mark style="color:green;">**not designed to act as a decision maker**</mark>, but as an idea generator and aid in rapid decision making.

### <mark style="color:purple;">How does it work?</mark>

IM-GPT's functionality hinges on its ability to <mark style="color:yellow;">tailor its analysis to the user's specified area of interest or content.</mark>&#x20;

Users have the flexibility to define their focus, whether it be industry sectors, geographic regions, investment themes, or individual securities.&#x20;

Additionally, users can provide specific information for analysis and instruct IM-GPT on the desired approach to scrutinise the data.&#x20;

### <mark style="color:purple;">Highlighting insights and anomalies</mark>

IM-GPT delivers comprehensive analyses encompassing:

1. A summary of the content
2. Historical context
3. Potential insights challenging existing perceptions
4. Implications for the future
5. Suggestions for further investigation
6. Potential courses of action

### <mark style="color:purple;">How has it been trained?</mark>

<mark style="color:green;">**Historical Datasets**</mark>

IM-GPT has access to the following datasets, accessed via retrieval augmented generation (RAG) and vector databases:&#x20;

·         Annual Reports

·         Company Announcements

·         Shareholder Register

·         Websites

·         Social Media

·         News Media

·         Academic Papers

·         Finance textbooks

<mark style="color:green;">**Training Dataset**</mark>

IM-GPT is a fine tuned version of several open sourced large language models.  It's fine tuning dataset encompasses the full range of investment management domains

·         Behavioural Economics

·         Industry Analysis

·         Valuation techniques

·         Systems Thinking

·         Organisational Management

·         Marketing and Distribution Theory

·         Theory of the Firm

·         Portfolio Theory

·         Quantitative Techniques

·         Balance Sheet

·         Cashflow Statements

·         ESG

·         Remuneration Analysis and Incentives

·         Value and Growth Investing

·         Risk and Volatility

### <mark style="color:purple;">In Action</mark>

The model is able to ingest streaming fundamental data - for example company announcements - and interpret them through the lens of fundamental analysis.

As important, it has <mark style="color:blue;">**access to history**</mark> - it can compare incoming fundamental data with historical data to detect anomalies.  This historical data is not just constrained to the individual company, but its competitors, suppliers, macroeconomic influences and any field related to the incremental news flow.

This enables IM-GPT to highlight to the user discrepancies between prevailing perceptions and reality.&#x20;

These insights, characterised by differing from consensus views and being accurate, are vital in identifying market inefficiencies.&#x20;

### <mark style="color:purple;">Psychological Biases</mark>

Behavioural economics explores how psychological factors affect economic decision-making.&#x20;

Within this field, heuristic biases describe mental shortcuts that people often use, sometimes leading to systematic errors or irrational behavior.&#x20;

IM-GPT has been trained on the vast domain of behavioural finance.  It is aware of the many areas of bias and error and seeks to correct them.

For example:

<mark style="color:green;">Anchoring Bias:</mark> Relying too heavily on the first piece of information encountered (the "anchor") when making decisions.

<mark style="color:green;">Availability Heuristic:</mark>  Making judgments about the likelihood of events based on how easily examples come to mind.

<mark style="color:green;">Representativeness Heuristic:</mark>  Assessing the similarity of objects and organizing them based on the category prototype.

<mark style="color:green;">Confirmation Bias:</mark> Tendency to search for, interpret, and remember information in a way that confirms one's pre-existing beliefs or values.

<mark style="color:green;">Overconfidence Bias:</mark> Overestimating one's abilities or the accuracy of one's beliefs, often leading to excessive risk-taking.

<mark style="color:green;">Hindsight Bias:</mark> The tendency to believe, after an event has occurred, that one would have predicted or expected it, commonly known as the "I-knew-it-all-along" effect.

<mark style="color:green;">Sunk Cost Fallacy:</mark> Continuing a behavior or endeavour based on previously invested resources (time, money, effort), even when it's not in the best interest.

<mark style="color:green;">Endowment Effect:</mark>  Valuing something more when you own it, leading to potentially irrational decisions like refusing to sell an asset for its market value.

<mark style="color:green;">Status Quo Bias:</mark> Favouring the current situation and resisting change, even when change might lead to a better outcome.

### <mark style="color:purple;">The product</mark>

IM-GPT is designed to give edge to traders and investors.  This edge has often been the domain of quantitative based investors and high frequency traders.  Fundamental investors now have their own tool to generate market edge.


# Employment Law

Navigating the intricaties of Australian employment law

### <mark style="color:purple;">Employment Legal Assistant</mark>

The Aleph-1 Employment Legal Assistant has been developed to assist both employers and employees navigate the highly complex and risky arena of employment law.

It has access to all relevant areas of employment law, every relevant act. &#x20;

It also has direct access to all Awards that outline the minimum pay rates and conditions of employment.  There are more than 100 industry or occupation awards that cover most people who work in Australia.

***

<table><thead><tr><th width="258">Factor</th><th></th></tr></thead><tbody><tr><td>Domain</td><td><strong>Professional Services</strong></td></tr><tr><td>Name</td><td>Aleph-1 Employment Law Assistant</td></tr><tr><td>Expertise</td><td>Employment Law and Awards Expert</td></tr><tr><td>Personality</td><td>Professional and Conversational</td></tr><tr><td>Training Data</td><td>Employment Law</td></tr><tr><td>Base Model</td><td>Llama3-13b</td></tr><tr><td>Training Platform</td><td>Continuum-Axolotl</td></tr><tr><td>Inference</td><td>Cloud L40S GPU</td></tr><tr><td>Data Privacy</td><td>Nemo Guardrails</td></tr><tr><td><mark style="color:green;">Knowledge Base</mark></td><td>Equal Opportunities Act 2010 </td></tr><tr><td></td><td>Age Discrimination Act 2004 </td></tr><tr><td></td><td>Fair Work Act 2009 </td></tr><tr><td></td><td>Racial Discrimination Act 2009 </td></tr><tr><td></td><td>Racial and Religious Toleration Act 2009 </td></tr><tr><td></td><td>Disability Discrimination Act 1992 </td></tr><tr><td></td><td>Occupational Health and Safety Act </td></tr><tr><td></td><td>Industrial Relations Act 1996 </td></tr><tr><td></td><td>Disability Discrimination Act 1992 </td></tr><tr><td></td><td>Sex Discrimination Act 1984</td></tr></tbody></table>

The Aleph-1 Employment Lawyer can be deployed as part of a standalone application, implemented within an existing knowledge base platform or customised depending on client requirements.

The model is not designed to replace qualified employment lawyer, but to provide guidance and advice that can be used to provide directions for next steps.

### <mark style="color:purple;">Problems Solved</mark>

Navigating Australia's employment laws is like embarking on a journey through a labyrinth of complexity.&#x20;

Here's a simplified breakdown of the issues faced by employees and employers alike.

<mark style="color:green;">**Extensive Legislation**</mark><mark style="color:green;">:</mark> The Fair Work Act 2009 is a beast of a document, covering everything from hiring to firing. Its size and complexity can leave business owners scratching their heads.

<mark style="color:green;">**Industry-Specific Awards**</mark><mark style="color:green;">:</mark> There are over 120 Modern Awards, each tailored to a specific industry. This means businesses must decipher which one applies to them, adding another layer of confusion.

<mark style="color:green;">**Small Business Struggles**</mark><mark style="color:green;">:</mark> Limited resources and legal know-how often leave small businesses feeling lost in the legal maze. Non-compliance can lead to headaches and penalties, diverting attention from business growth.

<mark style="color:green;">**Calls for Simplicity**</mark><mark style="color:green;">:</mark> There's a growing chorus calling for simpler laws and more support for small businesses. Streamlining legislation and offering tailored guidance could ease the burden, but this is not on the horizon.

<mark style="color:green;">**Constant Changes**</mark><mark style="color:green;">:</mark> Updates to laws and awards can feel like trying to hit a moving target. Staying up-to-date is a challenge, especially for those already juggling multiple tasks.

<mark style="color:green;">**Impact vs. Perception**</mark><mark style="color:green;">:</mark> While laws aim to protect workers, they can sometimes stifle business productivity. Finding the right balance is key.

<mark style="color:green;">**Tailored Support Needed**</mark><mark style="color:green;">:</mark> Small businesses crave resources that speak their language. Simplified guidelines and access to legal advice could make a world of difference.  Our employment law model can assist.

Australian employment laws are like a puzzle with too many pieces.  An AI driven employment employment law assistant delivers simplification and support to enable navigation of this complex arena.


# Psychology and Mental Health

Providing assistance to professional psychologists and mental health workers

### <mark style="color:purple;">Psychologist's Assistant</mark>

The Aleph-1 Psychologists Assistant has been developed to assist psychologists and mental health workers do their difficult jobs.

This model has been trained on gain a deep understanding of human behaviour, mental health disorders, therapeutic techniques, and ethical considerations.&#x20;

The psychology industry in Australia confronts several challenges that significantly impact the delivery and effectiveness of mental health services.&#x20;

One of the most pressing difficulties is ensuring that psychological services are equitably accessible across Australia, especially in rural and remote areas.&#x20;

These regions frequently suffer from a scarcity of mental health professionals, which prolongs wait times and diminishes the availability of care. The disparity in access exacerbates mental health outcomes for individuals living in these areas.

The use of an neural language model as an assistant in working and treating patients will increase productivity and result in better outcomes for patients.

This model can be customised for specific use cases, and integrated with patient databases in a fully secure environment.

The integration of technology and telehealth offers opportunities to improve access to psychological services. However, it also introduces challenges related to data security, privacy, and the effectiveness of remote therapy compared to in-person sessions.

***

<table><thead><tr><th width="258">Factor</th><th></th></tr></thead><tbody><tr><td>Domain</td><td><strong>Professional Services</strong></td></tr><tr><td>Name</td><td>Aleph-1 Psychologist and Mental Health Assistant</td></tr><tr><td>Expertise</td><td>Psychology and Mental Health</td></tr><tr><td>Personality</td><td>Professional and Interactive</td></tr><tr><td>Training Data</td><td>Psychology and Mental Health Academic Research</td></tr><tr><td>Base Model</td><td>Llama3-13b</td></tr><tr><td>Training Platform</td><td>Continuum-Axolotl</td></tr><tr><td>Inference</td><td>Cloud L40S GPU</td></tr><tr><td>Data Privacy</td><td>Nemo Guard </td></tr><tr><td><mark style="color:green;">Knowledge Base</mark></td><td>Cognitive-Behavioural Therapy (CBT)</td></tr><tr><td></td><td>Psychodynamic Therapy</td></tr><tr><td></td><td>Humanistic and Person-Centred Therapy</td></tr><tr><td></td><td>Family Systems Therapy</td></tr><tr><td></td><td>Existential Therapy</td></tr><tr><td></td><td>Sigmund Freud and texts</td></tr><tr><td></td><td>Skinner and texts</td></tr><tr><td></td><td>Jean Piaget and text</td></tr><tr><td></td><td>APS Code of Ethics</td></tr><tr><td></td><td>Australian Counselling Association Code of Ethics</td></tr></tbody></table>

### <mark style="color:purple;">Problems Solved</mark>

#### <mark style="color:green;">Funding and Resource Allocation</mark>

Funding for mental health services, including psychology, is often characterised by uneven distribution and insufficiency.  This financial constraint hinders the capacity of mental health providers to offer comprehensive care, particularly for individuals requiring complex or prolonged treatment.&#x20;

The challenge lies in securing adequate funding and efficiently allocating resources to areas of highest need without compromising the quality of care.

#### <mark style="color:green;">Navigating Health Insurance and Medicare</mark>

Psychologists and their clients frequently navigate the complexities of health insurance and Medicare, which can be daunting.&#x20;

Issues such as understanding coverage limits, rebate limitations, and the intricacies of the claim process can create barriers to accessing necessary psychological services.&#x20;

Simplifying these processes and increasing transparency could significantly improve access to mental health care.

#### <mark style="color:green;">Regulatory Compliance</mark>

Psychologists in Australia must adhere to stringent professional and ethical standards.&#x20;

Keeping abreast of regulatory changes and ensuring compliance demand continuous attention and effort, adding another layer of complexity.

#### <mark style="color:green;">Professional Development and Training</mark>

The field of psychology is dynamic, with constant developments in research, techniques, and best practices.&#x20;

Psychologists must engage in ongoing professional development, which can be challenging due to time constraints and the rapid pace of advancements in the field.

#### <mark style="color:green;">Integration with Other Health Services</mark>

Effective collaboration and integration with other health professionals are essential for providing holistic care.&#x20;

However, this necessitates efficient communication and coordination among various health services, which can be complex and time-consuming.

#### <mark style="color:green;">Workforce Challenges</mark>

The sustainability of the psychology workforce is a significant concern, with issues related to recruitment and retention, especially in rural and less urbanised areas.&#x20;

Addressing these challenges requires innovative approaches to attract and keep mental health professionals in these critical regions.

#### <mark style="color:green;">Diversity and Cultural Competency</mark>

Psychologists must be equipped to meet the diverse needs of their clients, accounting for cultural, linguistic, and demographic differences.&#x20;

Developing cultural competency and tailoring approaches to diverse populations are essential for delivering effective psychological care.

Generative AI model, trained in this field can address many of these challenges - enhancing the accessibility, quality, and effectiveness of psychological services across Australia.&#x20;


# Home Insurance

Providing assistance to consumers navigating the complexity of home insurance products

### <mark style="color:purple;">Home Insurance Expert</mark>

The Aleph-1 Home Insurance Expert has been developed to assist consumers navigate the complex landscape of home insurance in Australia

Apart from complex product design, risk is changing due to extreme weather and climate related risk - making it difficult for some home owners to gain insurance and for insurers to price risk.

***

<table><thead><tr><th width="258">Factor</th><th></th></tr></thead><tbody><tr><td>Domain</td><td><strong>Professional Services</strong></td></tr><tr><td>Name</td><td>Aleph-1 Home Insurance</td></tr><tr><td>Expertise</td><td>Insurance and Risk</td></tr><tr><td>Personality</td><td>Professional and Interactive</td></tr><tr><td>Training Data</td><td>Insurance Industry Data</td></tr><tr><td>Base Model</td><td>Llama3-13b</td></tr><tr><td>Training Platform</td><td>Continuum-Axolotl</td></tr><tr><td>Inference</td><td>Cloud L40S GPU</td></tr><tr><td>Data Privacy</td><td>Nemo Guard </td></tr><tr><td><mark style="color:green;">Knowledge Base</mark></td><td>General Insurance Code of Practice</td></tr><tr><td></td><td>Product Disclosure Statements</td></tr><tr><td></td><td>Australian Securities and Investment Commission</td></tr><tr><td></td><td>Policy Documents</td></tr><tr><td></td><td>Consumer Feedback (proprietary)</td></tr></tbody></table>

### <mark style="color:purple;">Problems Solved</mark>

<mark style="color:green;">Complex Product Design</mark>

A primary concern is the inherent complexity of home and contents insurance policies.&#x20;

This complexity often results in difficulty for consumers to effectively compare options, leading to <mark style="color:yellow;">unintentional underinsurance</mark>. Such policies, with their intricate details and fine print, can obscure the full extent of coverage, leaving many homeowners unaware of their vulnerability until it is too late.

#### <mark style="color:green;">Unaffordable Premiums</mark>

Alarmingly, a vast majority of policyholders have reported increases in their insurance premiums, highlighting a trend towards unaffordability, particularly in regions prone to natural disasters.&#x20;

This trend has disproportionately affected low-income households, many of which find themselves priced out of the insurance market altogether.&#x20;

The escalation in premiums, with an average rise of 28% within a year and spikes up to 50% in high-risk areas, underscores the urgent need for intervention to address affordability and access.

#### <mark style="color:green;">Inaccessible Risk Information</mark>

The challenge is further compounded by the difficulty in accessing reliable and comprehensive information on natural hazard risks.&#x20;

The scattered and often inaccurate data complicates the assessment of risks, hindering homeowners' ability to make informed decisions about insurance and risk mitigation.

#### <mark style="color:green;">Ignored Mitigation Efforts by Homeowners</mark>

Many homeowners are willing to invest in risk mitigation measures to potentially lower their insurance costs. However, these efforts are frequently overlooked by insurers when pricing policies, discouraging proactive measures that could benefit both the homeowner and the insurer in the long term.


# Consumer Surveying

A deeper and more nuanced interaction with consumers

### <mark style="color:purple;">Consumer Profiling</mark>

The Aleph-1 Home Consumer Profiling Model has been developed to assist organisations and marketing agencies engage with their customers to ask natural language questions to discover more nuanced views and uncover perceptions around their brand.

These interactions are supplemented by access to social media platforms.

These models can be fine tuned to act within survey design parameters and take on the persona of a partical brand and to engage empathatically and conversationally with consumers.

***

<table><thead><tr><th width="258">Factor</th><th></th></tr></thead><tbody><tr><td>Domain</td><td><strong>Marketing</strong></td></tr><tr><td>Name</td><td>Aleph-1 Consumer Profiling</td></tr><tr><td>Expertise</td><td>Brand Perception and Marketing</td></tr><tr><td>Personality</td><td>Friendly and conversational</td></tr><tr><td>Training Data</td><td>Focus groups and Historical Consumer Interactions</td></tr><tr><td>Base Model</td><td>Llama3-13b</td></tr><tr><td>Training Platform</td><td>Continuum-Axolotl</td></tr><tr><td>Inference</td><td>Cloud L40S GPU</td></tr><tr><td>Data Privacy</td><td>Nemo Guard </td></tr><tr><td><mark style="color:green;">Knowledge Base</mark></td><td>Interview Techniques</td></tr><tr><td></td><td>Foundations of Research Methodology</td></tr><tr><td></td><td>Brand Equity and Brand Management</td></tr><tr><td></td><td>Facilitation Techniques</td></tr></tbody></table>

### <mark style="color:purple;">Problems Solved</mark>

#### <mark style="color:green;">Streamlining Survey Collection</mark>

* <mark style="color:purple;">**Efficiency:**</mark> The model automates and expedites the survey collection process, allowing businesses to gather consumer feedback quickly. This rapid collection enables companies to stay agile, responding to consumer needs and market changes in a timely fashion.

#### <mark style="color:green;">Personalising Surveys</mark>

* <mark style="color:purple;">Customisation</mark>: By personalising surveys based on individual consumer histories or preferences, models increase engagement rates.  Personalised surveys are more likely to elicit honest and detailed feedback, providing richer insights into brand perception.

#### <mark style="color:green;">Sentiment Analysis</mark>

* <mark style="color:purple;">**Targeted Improvements**</mark><mark style="color:purple;">:</mark> The model has the ability to perform sentiment analysis, particularly on open-ended responses, offers precise insights into consumers' feelings about a brand. Identifying areas of satisfaction or concern allows for targeted actions to enhance the brand experience.

#### <mark style="color:green;">Predictive Analysis</mark>

* <mark style="color:purple;">**Future Insights:**</mark> Analysing trends in consumer feedback helps predict future behaviours and preferences. These predictions are crucial for guiding product development, tailoring marketing strategies, and making informed business decisions.

#### <mark style="color:green;">Enhancing Product Innovation</mark>

* <mark style="color:purple;">**Customer-Driven Innovation:**</mark> By analysing what consumers are saying about products or services, models can identify desired features or innovations. This customer-driven approach ensures that new or improved offerings align closely with consumer expectations.

#### <mark style="color:green;">Streamlining Market Research</mark>

* <mark style="color:purple;">**Dynamic Surveys:**</mark> your model can create surveys that adapt questions based on respondent answers, leading to deeper insights. This dynamic approach uncovers nuanced understandings of consumer behaviour and market trends.

#### <mark style="color:green;">Customer Sentiment Analysis</mark>

* <mark style="color:purple;">**Deeper Consumer Insights**</mark><mark style="color:purple;">:</mark> A model can offer a sophisticated analysis of qualitative feedback, revealing nuanced sentiments. Understanding these deeper consumer sentiments helps address not just the symptoms of dissatisfaction but its root causes.

#### <mark style="color:green;">Automating Follow-ups and Real-time Monitoring</mark>

* <mark style="color:purple;">**Prompt Action**</mark><mark style="color:purple;">:</mark> Automated follow-ups and real-time monitoring of feedback ensure that businesses can react promptly to consumer needs. This level of responsiveness is key to maintaining customer satisfaction and loyalty.

The Continuum Consumer Profiling model transforms how brands gather, analyse, and act on consumer perceptions.&#x20;

This model not only makes the process more efficient and personalised but also ensures that insights are actionable, directly influencing product innovation, marketing strategy, and overall customer experience.

By leveraging the capabilities of our model, businesses can achieve a deeper, more nuanced understanding of their brand's perception in the marketplace.


# Government Grants

### <mark style="color:purple;">Government Grant Model</mark>

The Aleph-1 Government Grant Model has been developed to assist organisations become aware of and interact with the Federal and State government grant program.

***

<table><thead><tr><th width="258">Factor</th><th></th></tr></thead><tbody><tr><td>Domain</td><td><strong>Government Services</strong></td></tr><tr><td>Name</td><td>Aleph-1 Grant Application Model</td></tr><tr><td>Expertise</td><td>Government Grants and Tender Process</td></tr><tr><td>Personality</td><td>Professional and Interactive</td></tr><tr><td>Training Data</td><td>Insurance Industry Data</td></tr><tr><td>Base Model</td><td>Mistral-13b</td></tr><tr><td>Training Platform</td><td>Continuum-Axolotl</td></tr><tr><td>Inference</td><td>Cloud L40S GPU</td></tr><tr><td>Data Privacy</td><td>Nemo Guard </td></tr><tr><td><mark style="color:green;">Knowledge Base</mark></td><td>Commonwealth Grants Rules and Guidelines 2017 (CGRGs)</td></tr><tr><td></td><td>The Public Governance, Performance and Accountability Act 2013 (PGPA Act)</td></tr><tr><td></td><td>Resource Management Guide No. 411 Grants, Procurements or other financial arrangement (RMG 411)</td></tr><tr><td></td><td>Policy Documents</td></tr></tbody></table>

### <mark style="color:purple;">Problems Solved</mark>

The Continuum Grant Assistant Model can significantly streamline and enhance the complex process of grant application and administration for organisations.&#x20;

Here's how it addresses the various challenges and complexities involved:

#### <mark style="color:green;">Determining the Type of Funding Arrangement</mark>

* <mark style="color:purple;">**Automated Guidance**</mark><mark style="color:purple;">:</mark> The model can automatically guide organisations through the process of determining the nature of their funding needs using tools like the Grants Decision Tree, ensuring that they pursue the correct type of financial support.

#### <mark style="color:green;">Compliance with the Grants Framework and Finance Law</mark>

* <mark style="color:purple;">**Regulatory Alignment**</mark><mark style="color:purple;">:</mark> By being fine-tuned to understand and interpret the Commonwealth Grants Rules and Guidelines 2017 (CGRGs), the model ensures that all grant designs and deliveries are in strict compliance with regulatory requirements, minimizing legal risks.

#### <mark style="color:green;">Complex Product Design and Approval Process</mark>

* <mark style="color:purple;">**Streamlined Planning**</mark><mark style="color:purple;">:</mark> The Grant Assistant Model can assist in the planning and design stages of grant opportunities, integrating considerations like stakeholder consultation, risk management, and governance arrangements, thereby simplifying the approval process.

#### <mark style="color:green;">Developing Grant Opportunity Guidelines</mark>

* <mark style="color:purple;">**Guideline Development**</mark><mark style="color:purple;">:</mark> With access to government-provided templates and a deep understanding of compliance requirements, the model can help draft clear and authoritative grant opportunity guidelines, expediting the approval process.

#### <mark style="color:green;">Risk Assessment and Management</mark>

* <mark style="color:purple;">**Risk Analysis Support**</mark><mark style="color:purple;">:</mark> The model can aid in conducting thorough risk analyses of grant programs, facilitating consultations with relevant departments, and helping to agree on risk ratings to streamline the subsequent processes.

#### <mark style="color:green;">Streamlined Approval Process</mark>

* <mark style="color:purple;">**Efficiency in Approvals**</mark><mark style="color:purple;">:</mark> For grants identified as low-risk, the model can ensure a swift approval process by accurately using government templates. For medium to high-risk grants, it can help prepare the necessary documentation for ministerial approvals.

#### <mark style="color:green;">Use of Grants Administration Hubs</mark>

* <mark style="color:purple;">**Integration with Administration Hubs**</mark><mark style="color:purple;">:</mark> The model can serve as an interface between organisations and Grants Administration Hubs, offering guidance on design considerations and helping to navigate the administration landscape efficiently.

#### <mark style="color:green;">Consideration of Grant Applicants</mark>

* <mark style="color:purple;">**Applicant-Focused Design**</mark><mark style="color:purple;">:</mark> By simulating potential applicant interactions, the model can help ensure that grant programs are designed with a focus on minimising the application burden, making the process more accessible and less daunting for applicants.

#### <mark style="color:green;">Legal, Accounting, and Policy Matters</mark>

* <mark style="color:purple;">**Comprehensive Governance**</mark><mark style="color:purple;">:</mark> The model can assist in addressing legal, accounting, and policy considerations during the development of grant programs, ensuring robust governance arrangements are in place.

#### <mark style="color:green;">Streamlining Government Grants Administration</mark>

* <mark style="color:purple;">**Optimisation of Administration**</mark><mark style="color:purple;">:</mark> Through its ability to process and analyse large volumes of information, the model can contribute to government efforts to consolidate grants administration, enhancing expertise and improving the applicant experience.

#### <mark style="color:green;">Exemption and Deferral Requests</mark>

* <mark style="color:purple;">**Handling Exceptions**</mark><mark style="color:purple;">:</mark> The model can help manage the process of seeking exemptions or deferrals, providing the necessary documentation and rationale to support such requests.

<mark style="color:green;">**Early Engagement and Optimisation**</mark>

The model facilitates early engagement with Grants Administration Hubs, offering insights into service delivery quotes and optimising the grant design process based on the specific needs of in-scope entities.

The Continuum Grant Assistant Model acts as a comprehensive assistant throughout the grant application and administration process.&#x20;

By automating routine tasks, providing regulatory guidance, streamlining the approval process, and focusing on the needs of applicants, it greatly reduces the complexity and enhances the efficiency of dealing with grants, making the entire process more manageable for organisations.


# Aged Care

Customised Fine-Tuned Large Language Models in the Aged Care Sector: A Comprehensive Use Case

The aged care sector in Australia faces a multitude of challenges ranging from workforce development, governmental policy alignment, service accessibility, to personalising care for the elderly.

AI presents a transformative solution to these issues, enhancing the efficiency, effectiveness, and personalisation of aged care services.

### <mark style="color:purple;">**Workforce Development**</mark>

Customised models can assist workforce training within the aged care sector.

By analysing individual caregiver performance, learning preferences, and professional gaps, AI-driven platforms can generate tailored educational content.

This includes interactive simulations and scenario-based training modules, offering a personalised learning experience.&#x20;

### <mark style="color:purple;">**Government Policy and Market Dynamics**</mark>

While directly influencing government policies or market structures may be beyond the scope of models, their capacity to analyse vast datasets for care delivery patterns, patient outcomes, and service utilisation can provide critical insights.

These insights can guide evidence-based recommendations for policy adjustments, aiming to improve care quality and accessibility.

### <mark style="color:purple;">**Personalisation in Care Coordination**</mark>

AI models stand to significantly ease the process of navigating the aged care system for seniors and their families.&#x20;

Acting as virtual assistants, these models can offer personalised advice, answer queries, and guide users through service applications based on individual care needs and preferences.&#x20;

For care coordinators, generative AI can be indispensable tools, providing updated information on service availability and regulatory standards to facilitate comprehensive and individualised care planning.

### <mark style="color:purple;">**Enhancing Information Accessibility and Support**</mark>

Models can drive sophisticated virtual assistants capable of providing instant and accurate responses to inquiries from care recipients, families, and caregivers.&#x20;

This immediate access to crucial information not only makes the aged care system more navigable but also allows human staff to concentrate on complex care tasks.

### <mark style="color:purple;">**Training and Education**</mark>

Through the generation of customised training materials, caregivers can be continuously educated on the best practices, latest research findings, and regulatory requirements, thereby elevating the standard of care provided.

### <mark style="color:purple;">**Policy and Research Analysis**</mark>

Models can efficiently summarise extensive research documents and policy papers, enabling administrators and policymakers to stay informed about the latest developments in the aged care sector, thus facilitating informed decision-making.

### <mark style="color:purple;">**Administrative Support**</mark>

By organising and interpreting large volumes of unstructured data, such as patient feedback and care records, generative AI can assist streamline administrative processes, identify trends, and inform decisions, enhancing operational efficiency indirectly.

### <mark style="color:purple;">**Communication Enhancements**</mark>

Generative AI can also aid in overcoming language barriers between care recipients, families, and caregivers, fostering a more inclusive care environment. Additionally, they can create personalised content and engage in meaningful conversations, offering companionship and basic mental health support to older adults.


# Pharmaceuticals Benefit Scheme

The Australian healthcare system is characterised by a complex arrangement that divides funding and service provision between federal and state governments, leading to issues like service gaps and inefficiency.&#x20;

Only a small fraction of healthcare spending goes towards preventive measures. This is problematic given that a significant portion of chronic disease is preventable.&#x20;

Recommendations include considering alternative funding models to foster more integrated, patient-centred care focusing on preventive health.

The system has not equally benefitted Indigenous and non-Indigenous Australians.&#x20;

Aboriginal Community Controlled Health Organisations (ACCHOs) have been effective in delivering primary healthcare, and initiatives like the Special Pharmaceutical Benefits Scheme Agreement and Closing the Gap PBS strategy have improved medication access for Indigenous people.

### <mark style="color:purple;">**Challenges with the Pharmaceutical Benefits Scheme (PBS)**</mark>

Generic drugs in Australia are more expensive compared to other countries. Policies to reduce PBS costs include mandatory price reductions and establishing two drug formularies, which have achieved cost savings.

### <mark style="color:purple;">**Issue of Out-of-Pocket Costs**</mark>

Rising out-of-pocket costs are a major concern, with healthcare expenses increasing faster than the consumer price index. These costs, higher for people with chronic conditions and those in disadvantaged groups, often lead to deferred or forgone healthcare.

### <mark style="color:purple;">**Inequity in Healthcare Access**</mark>

Despite Medicare's aim of equitable healthcare access, disparities exist based on socio-economic status, geography, and in the distribution of hospital-based care. Rural and outer suburban residents face significant access challenges.


# Three ideas for autonomous agent applications

The research on autonomous agents has been a long-standing focus in both academia and industry, with the ultimate goal of achieving artificial general intelligence (AGI).&#x20;

Traditional approaches to developing autonomous agents often involve training them with limited knowledge within isolated environments, which differs significantly from how humans learn and makes it challenging for these agents to make human-like decisions.&#x20;

However, the recent advent of large language models (LLMs) has shown promising results in attaining human-level intelligence, sparking a surge of interest in <mark style="color:blue;">**LLM-based autonomous agents.**</mark>

In this <mark style="color:blue;">**April 2024**</mark> paper, the authors present a comprehensive survey of studies on LLM-based autonomous agents, providing a systematic review from a holistic perspective.&#x20;

The authors propose a *<mark style="color:yellow;">**unified framework**</mark>* that encompasses most of the previous work on designing agent architectures to better leverage LLMs.

{% embed url="<https://arxiv.org/abs/2308.11432>" %}
A Survey on Large Language Model based Autonomous Agents
{% endembed %}

<figure><img src="/files/OF9QFvdw3OhrNLOOQrN1" alt=""><figcaption><p>Aunified framework for the architecture design of LLM-based autonomous agent</p></figcaption></figure>

### <mark style="color:purple;">Drawing on inspiration from this highly cited paper on autonomous AI agents</mark>

### <mark style="color:blue;">**Personalised Mental Health Companion**</mark>&#x20;

Leveraging the profiling and memory modules, an LM-based autonomous agent could be designed to serve as a personalised mental health companion.&#x20;

The agent would gather information about the user's demographic background, personality traits, and mental health history to create a unique profile.&#x20;

As the user interacts with the agent over time, it would store their conversations, emotions, and coping strategies in its memory module.&#x20;

The <mark style="color:blue;">**planning module**</mark> would enable the agent to provide tailored support and guidance based on the user's specific needs and past experiences.&#x20;

The <mark style="color:blue;">**action module**</mark> would focus on engaging in empathetic communication and offering evidence-based strategies for managing stress, anxiety, and other mental health challenges.&#x20;

This application combines elements from psychology, social science, and natural language processing to create an accessible and adaptive mental health support system.

### <mark style="color:blue;">Collaborative Scientific Research Platform</mark>&#x20;

Drawing inspiration from the natural science applications discussed in the paper, a collaborative scientific research platform could be developed using LM-based autonomous agents.&#x20;

The platform would consist of multiple agents with specialised capabilities, such as literature review, hypothesis generation, experiment design, data analysis, and scientific writing.&#x20;

Researchers would interact with the platform using natural language, describing their research questions and objectives.&#x20;

The agents would then work together, leveraging their individual strengths and the collective knowledge stored in their memory modules.&#x20;

The <mark style="color:blue;">planning module</mark> would enable the agents to break down complex research tasks into manageable steps and adapt their strategies based on feedback from the researchers and the results of experiments.&#x20;

The <mark style="color:blue;">**action module**</mark> would focus on generating human-readable outputs, such as literature summaries, experiment protocols, data visualisations, and draft manuscripts.&#x20;

This application combines elements from documentation and data management, experiment assistance, and natural science education to create a powerful tool for accelerating scientific discovery.

### <mark style="color:blue;">Intelligent Urban Planning Assistant</mark>&#x20;

Combining ideas from civil engineering and social simulation, an intelligent urban planning assistant could be developed using LLM-based autonomous agents.&#x20;

The agent would be trained on a diverse dataset of urban planning projects, including information about demographics, infrastructure, transportation, and sustainability.&#x20;

The <mark style="color:blue;">**profiling module**</mark> would enable the agent to understand the unique characteristics and challenges of a given city or neighbourhood.&#x20;

The <mark style="color:blue;">**memory module**</mark> would store relevant case studies, best practices, and stakeholder feedback.&#x20;

The <mark style="color:blue;">**planning module**</mark> would generate multiple scenarios and evaluate their potential impacts using simulations that model complex social and economic dynamics.&#x20;

The <mark style="color:blue;">**action module**</mark> would focus on presenting clear, actionable recommendations to urban planners and policymakers, along with visualizations and interactive tools for exploring different options.&#x20;

This application combines elements from civil engineering, social simulation, and data visualization to create a data-driven and stakeholder-centric approach to urban planning.

These three applications demonstrate how the ideas presented in the paper can be creatively combined and adapted to solve complex problems in mental health, scientific research, and urban planning.&#x20;

By leveraging the strengths of LLM-based autonomous agents and drawing upon insights from multiple domains, these applications have the potential to create significant value and drive innovation in their respective fields.


# Financial Statement analysis with large language models

The University of Chicago, Booth School of Business

#### <mark style="color:green;">Research design</mark>

The authors provide standardised and anonymised financial statements (balance sheet and income statement) to GPT-4 and instruct the model to analyse them to determine the direction of future earnings. *<mark style="color:yellow;">No narrative or industry-specific information is provided</mark>*.

#### <mark style="color:green;">Comparison with human analysts</mark>

GPT-4's performance is compared to that of human financial analysts.&#x20;

The authors find that when using a [<mark style="color:blue;">**chain-of-thought (CoT) prompt**</mark>](/agents/what-is-agency/ai-reasoning-a-deep-dive-into-chain-of-thought-prompting) to emulate human reasoning, GPT-4 achieves a 60% accuracy in predicting the direction of future earnings, outperforming the median financial analyst.

#### <mark style="color:green;">Strengths and weaknesses of LLM vs. human analysts</mark>

While human analysts rely on soft information and broader context not available to the model, GPT-4's insights are more valuable when humans struggle with forecasts or when human forecasts are prone to biases or inefficiency.

#### <mark style="color:green;">Comparison with specialized ML models</mark>

GPT-4's accuracy is on par with or slightly higher than state-of-the-art machine learning models, such as logistic regression and artificial neural networks (ANNs), specifically trained for earnings prediction. GPT-4 and ANNs are found to be complementary, with GPT-4 performing well when ANNs struggle, especially for small or loss-making companies.

<mark style="color:blue;">**Reasons for GPT-4's success**</mark>

The authors *<mark style="color:yellow;">rule out the hypothesis that GPT-4's performance is driven by its memory</mark>*.  Instead, they find that GPT-4 generates useful narrative insights, such as ratio analysis, which are informative about future performance. These narratives, derived from CoT reasoning, are responsible for the model's superior performance.

<mark style="color:blue;">**Economic usefulness**</mark>

The authors demonstrate the economic usefulness of GPT-4's forecasts by analysing their value in predicting stock price movements.  Long-short strategies based on GPT-4 forecasts outperform the market and generate significant alphas and Sharpe ratios, particularly for small companies.

### <mark style="color:purple;">Analysis</mark>

#### <mark style="color:green;">Implications for financial analysis</mark>

The study suggests that LLMs like GPT-4 can play a role in financial decision-making, potentially transforming the way financial statement analysis is performed, and earnings forecasts developed.

The ability of LLMs to *<mark style="color:yellow;">**generate valuable insights without relying on narrative context**</mark>* highlights their potential to complement and even outperform human analysts.

#### <mark style="color:green;">Limits of LLMs</mark>

The paper provides evidence on the ability of LLMs to excel in quantitative tasks that require intuition and human-like reasoning, extending their capabilities beyond their native textual domain.&#x20;

This points towards the emergence of Artificial General Intelligence and suggests that the boundaries of LLMs are broader than previously thought.  I think this is a stretch, but that is their view.

{% file src="/files/y4Ik3cXqAhXUkg5eXNEd" %}
The paper
{% endfile %}

### <mark style="color:purple;">Conceptual Underpinnings</mark>

#### <mark style="color:green;">Financial analysts' approach to earnings forecasting</mark>

* Analysts begin with a systematic analysis of financial statements, often using standardised templates for consistency and accuracy.
* They establish a baseline understanding of a company's financial position and performance by assessing factors such as operating performance and capital structure.
* Analysts then contextualise the financial data by drawing upon their industry knowledge and private information about the firm before issuing forecasts.  Their objective is to accurately forecast company earnings.

#### <mark style="color:green;">Limitations of human analysts</mark>

* Despite generally outperforming time series models in producing credible annual earnings forecasts, financial analysts are naturally prone to errors and biases.
* Analysts may make <mark style="color:yellow;">technical errors</mark>, <mark style="color:yellow;">questionable economic judgments</mark>, or <mark style="color:yellow;">overreact to recent events</mark>, highlighting the complexity of processing large volumes of data efficiently.

#### <mark style="color:green;">Potential of LLMs in financial statement analysis</mark>

* General-purpose language models, such as ChatGPT, hold promise in facilitating financial statement analysis and associated tasks like earnings forecasting and decision-making.
* LLMs are noted for their knowledge across various domains and ability to quickly and efficiently process large quantities of data.
* They have demonstrated proficiency in answering CFA or CPA exam questions, processing large sets of financial data, and predicting certain economic outcomes.

#### <mark style="color:green;">Challenges faced by LLMs in financial statement analysis</mark>

* Financial statement analysis is a broad task that <mark style="color:yellow;">requires common sense, intuition, reasoning, and judgment</mark>, whereas machines typically excel in narrow, well-defined tasks.
* LLMs are not specifically trained to analyse financial information and have struggled with understanding the numeric domain.
* Humans are more capable of incorporating their knowledge of broader context, such as soft information, industry knowledge, and regulatory, political, and macroeconomic factors.

#### <mark style="color:green;">Potential advantage of LLMs</mark>

* LLMs' training on a vast body of general knowledge, encompassing business cases, financial theories, and economic contexts, may allow them to infer insights even from unfamiliar data patterns.
* This broader theoretical foundation could provide an advantage in the complex domain of financial analysis, where human experience, intuition, and judgment are valuable.

The conceptual underpinnings section highlights the significance of financial statement analysis, the role of human analysts, and the potential challenges and opportunities for LLMs in this domain.

It sets the stage for the paper's investigation into whether an LLM can successfully perform financial statement analysis tasks at a level comparable to professional human analysts, despite the inherent challenges posed by the complex and judgment-based nature of the task.

### <mark style="color:purple;">Methodology and Data</mark>

The methodology and data section of the paper provides a detailed explanation of how the authors use a large language model (LLM), specifically GPT-4, to analyse financial statements and predict earnings changes.&#x20;

#### <mark style="color:green;">Earnings prediction task</mark>

* Earnings prediction is a complex task that combines qualitative and quantitative analyses and involves professional judgment.
* The authors model how analysts make earnings predictions using a chain-of-thought (CoT) prompt with GPT-4.
* They focus on a relatively narrow information set that includes numerical information reported on the face of two primary financial statements (<mark style="color:yellow;">balance sheet and income statement</mark>), without textual information or broader context.
* This approach allows them to test the limits of the model when analysing financials and deriving insights from numeric data, which LLMs are not designed or trained to do.

<mark style="color:green;">**Prompts for financial statement analysis (FSA) and earnings prediction: a. "Simple" prompt**</mark>

* Instructs the LLM to analyse the two financial statements of a company and determine the direction of future earnings.
* Does not provide further guidance on how to approach the prediction task.

#### <mark style="color:blue;">Chain-of-Thought (CoT) prompt</mark>

Breaks down the problem into steps that parallel those followed by human analysts, effectively ingraining the methodology into the model and guiding it to mimic human-like reasoning.

* Instructs the model to take on the role of a financial analyst and perform financial statement analysis by:

&#x20;i. Identifying *<mark style="color:yellow;">**notable changes in certain financial statement items**</mark>*.&#x20;

ii. *<mark style="color:yellow;">Computing key financial ratios</mark>* without explicitly limiting the set of ratios.

iii. Providing *<mark style="color:yellow;">**economic interpretations**</mark>* of the computed ratios.

Based on the quantitative information and insights, the model is instructed to predict whether earnings are likely to increase or decrease in the subsequent period and produce a paragraph elaborating its rationale.

The model is also prompted to provide the predicted magnitude of earnings change (large, moderate, or small) and the confidence in its answer (ranging from zero to one).

<mark style="color:green;">**GPT-4 configuration**</mark>

* The authors use gpt-4-0125-preview, the most updated GPT model by OpenAI at the time of their experiment.
* <mark style="color:yellow;">**Temperature parameter is set to zero to ensure minimal variability**</mark> in the model's responses.
* Max tokens are not specified, and the top-p sampling parameter is set to one.
* The logprobs option is enabled to obtain token-level logistic probability values.

#### <mark style="color:blue;">Data:</mark>

<mark style="color:purple;">Compustat annual financial data (1968-2021)</mark>

* The authors use the entire universe of Compustat annual financial data from 1968 to 2021 fiscal years.
* They set aside data for 2022 to predict 2023 fiscal year earnings to test the robustness of the model's performance outside GPT's training window (ending in April 2023).
* Filters are applied to ensure data quality and consistency, *<mark style="color:yellow;">**resulting in 150,678 observations from 15,401 distinct firms**</mark>*.
* For each firm-year, the balance sheet and income statement are reconstructed using Compustat data, following Capital IQ's balancing model, and any identifying information is omitted.

<mark style="color:purple;">BES data (1983-2021):</mark>

* For the analysis involving analyst forecasts, the authors use data from IBES, starting the sample in 1983.
* Individual forecasts are extracted, and monthly consensus forecasts are constructed.
* The sample is restricted to firm-years with at least three analyst forecasts, resulting in 39,533 firm-year observations.

#### <mark style="color:green;">**Descriptive statistics**</mark>

* Panel A describes the full sample (1968-2021), revealing that approximately 55.5% of observations report an actual increase in earnings (Target), while GPT prediction (Pred GPT) implies an average of 53.0% of observations will experience an earnings increase.
* Panel B is restricted to the analyst sample (1983-2021) and includes analyst forecasts issued within one, three, and six months from the previous year's earnings release.
* Compared to GPT, financial analysts tend to be slightly more pessimistic in their forecasts.
* Companies in the Analyst Sample are, on average, larger in size, have a lower book-to-market ratio, higher leverage, and lower earnings volatility compared to the full sample, but are similar in terms of the actual frequency of EPS increases.

The methodology and data section highlights the authors' approach to using GPT-4 for financial statement analysis and earnings prediction, focusing on a narrow information set to test the model's limits.&#x20;

The use of both simple and CoT prompts allows for a comparison of the model's performance with and without guided reasoning. The data from Compustat and IBES provides a comprehensive sample for testing the model's predictions and comparing them to human analysts' forecasts.

### <mark style="color:purple;">Performance versus the analysts</mark>

The paper compares the performance of GPT-4 in predicting the direction of future earnings based on financial statement analysis to that of financial analysts.&#x20;

The authors use several methods to evaluate the model's performance and explore the complementarity between human analysts and GPT. Here's a critical assessment of the main results:

#### <mark style="color:green;">Prediction accuracy</mark>

* GPT-4 with a simple prompt achieves an accuracy of 52.33% and an F1-score of 54.52%, which is on par with the first-month consensus forecasts by financial analysts following the earnings release.
* When using <mark style="color:yellow;">**chain-of-thought (CoT) prompts**</mark>, GPT-4 achieves an accuracy of 60.35%, *<mark style="color:yellow;">**outperforming analyst predictions by 7 percentage points**</mark>*, even without access to narrative or contextual information available to analysts.
* While the results are impressive, it's important to note that the paper does not provide a detailed explanation of how the CoT prompts were designed or optimised.&#x20;
* The specific instructions used in the CoT prompts could have a significant impact on the model's performance, and more transparency in this regard would strengthen the findings.

#### <mark style="color:green;">Complementarity between human analysts and GPT</mark>

The authors explore instances where forecasts are erroneous and find that *<mark style="color:yellow;">**GPT-4's predictions are more likely to be inaccurate for smaller firms**</mark>*, firms with higher leverage ratios, loss-making firms, and firms with volatile earnings.&#x20;

However, the magnitude of these effects is smaller for human analysts, suggesting that they *<mark style="color:yellow;">**benefit from access to soft information and additional context**</mark>*.

* The incremental informativeness analysis shows that both GPT-4 and analyst forecasts are positively associated with future outcomes, and their combined use improves the adjusted R-squared, indicating complementarity.
* While these findings are valuable, the paper does not look into the specific types of soft information or context that human analysts might rely on.  A more detailed discussion of these factors could provide a clearer understanding of the relative strengths and weaknesses of GPT-4 and human analysts.

Overall, the paper presents compelling evidence that GPT-4 can outperform human analysts in predicting the direction of future earnings based on financial statement analysis, even without access to the same level of contextual information.&#x20;

The authors also demonstrate the complementarity between GPT-4 and human analysts, highlighting the potential for the model to add value in situations where humans struggle.

However, the paper could benefit from more transparency in the design of the CoT prompts and a more detailed discussion of the specific factors that contribute to the relative strengths and weaknesses of GPT-4 and human analysts.&#x20;

Additionally, the authors could explore the potential limitations of GPT-4, such as its ability to handle novel or unusual financial situations that may not be well-represented in its training data.

Despite these limitations, the paper makes a significant contribution to the literature on the application of large language models in financial analysis and provides a strong foundation for future research in this area.

### <mark style="color:purple;">Predictive Ability</mark>

This section aims to understand the sources of GPT-4's predictive ability and explores two potential explanations:&#x20;

the model's memory and its ability to generate narrative insights based on numeric data. The authors attempt to rule out the possibility of look-ahead bias and investigate whether the model's generated texts are informative.&#x20;

Here are some potential biases and ways that could annul the experimental findings:

<mark style="color:green;">**Look-ahead bias**</mark>

* The authors argue that their research design is relatively immune to look-ahead bias because they use a *<mark style="color:yellow;">**consistent anonymised format for financial statements**</mark>*, making it difficult for the model to infer a firm's identity or the specific year.
* However, there might be subtle patterns or combinations of financial ratios that are unique to certain companies or industries, which GPT-4 could potentially recognise based on its training data. If this is the case, the model's predictive ability could be overstated.
* To further investigate this, the authors could conduct additional experiments with synthetic financial data that preserve the overall statistical properties of the original data but break any potential links to specific companies or industries.

#### <mark style="color:green;">Experiment design</mark>

* The authors use chain-of-thought (CoT) prompts to guide GPT-4 in analysing financial statements, which could inadvertently introduce bias in the model's predictions.
* The specific instructions provided in the CoT prompts might steer the model *<mark style="color:yellow;">**towards focusing on certain financial ratios or trends that are known to be predictive of future earnings**</mark>*, thus inflating its performance.  This would be good!
* To address this concern, the authors could experiment with different sets of CoT prompts that vary in their level of specificity and guidance to assess the sensitivity of the results to the prompt design.

<mark style="color:green;">**Selection bias**</mark>

* The paper uses the entire universe of Compustat annual financial data from 1968 to 2021, which could introduce selection bias if the dataset is not representative of the broader population of companies.
* If GPT-4's training data overrepresents certain types of companies or industries that are more predictable, the model's performance might be overstated.
* To mitigate this issue, the authors could conduct robustness tests using alternative datasets or by stratifying the sample based on company characteristics to ensure that the results are not driven by specific subsets of the data.

<mark style="color:green;">**Temporal bias**</mark>

* The authors find that GPT-4's predictive accuracy decreases over time, with *<mark style="color:yellow;">**sharp drops during international macroeconomic downturns**</mark>*.
* If the model's training data is skewed towards more recent years or if it overrepresents certain economic conditions, its predictive ability might not generalise well to other time periods or market environments.
* To address this concern, the authors could perform additional tests using rolling windows or by explicitly controlling for macroeconomic factors to assess the stability of the model's performance across different time periods.

<mark style="color:green;">**Information leakage**</mark>

* While the authors aim to use only numerical data from financial statements, there is a risk that the process of extracting and preprocessing the data could inadvertently introduce information leakage.
* For example, if the data preparation process involves any form of normalization or scaling based on future information, it could artificially inflate the model's predictive performance.
* To mitigate this risk, the authors should carefully review their data preprocessing pipeline and ensure that all transformations are based solely on information available at the time of prediction.

<mark style="color:green;">**Confounding factors**</mark>

* The authors argue that GPT-4's predictive ability stems from its capacity to generate narrative insights based on numeric data.
* However, there might be confounding factors that drive both the model's generated texts and its predictive performance, such as the underlying quality or complexity of the financial statements.
* To disentangle these effects, the authors could control for various measures of financial statement quality or complexity in their analyses and assess whether GPT-4's predictive ability persists after accounting for these factors.

While the authors have taken steps to address potential biases and alternative explanations, further tests and robustness checks could help strengthen the validity of their findings. By considering and addressing these potential issues, the paper can provide more convincing evidence of GPT-4's ability to generate valuable insights from financial statements and its potential for enhancing financial analysis.

### <mark style="color:purple;">Trading Strategy Performance</mark>

In this section, the authors investigate the practical value of GPT-4-based financial statement analysis by evaluating the performance of trading strategies based on the model's output.&#x20;

They argue that if GPT-4's forecasts contain incremental information about future profitability, they should also predict future stock returns.&#x20;

The authors compare the performance of three types of strategies: one based on GPT-4 forecasts, and two others based on artificial neural network (ANN) and logistic regression forecasts that rely on numeric information.&#x20;

Here's a detailed explanation of the methodology and results:

<mark style="color:green;">**Methodology:**</mark>

a. Portfolio formation:

* The authors form portfolios on June 30 of each year, allowing approximately three months for the market to process the reported financial information, and hold the portfolios for one year.
* For ANN and logistic regression strategies, stocks are sorted into ten portfolios based on the predicted probabilities of earnings increase. The strategies take long positions in the top decile stocks and short positions in the bottom decile.
* For GPT-4 strategies, the authors use binary directional predictions, magnitude predictions, and average log probabilities of tokens to form portfolios. They select stocks predicted to experience a "moderate" or "large" earnings increase, sort them by log probability values, and retain the top 10% with the highest expected confidence for long positions. Similarly, they select stocks predicted to experience a "moderate" or "large" earnings decrease, sort them by log probability values, and short the top 10% with the highest expected confidence.

b. Performance evaluation:

* The authors compute Sharpe ratios for equal-weighted and value-weighted portfolios, with monthly rebalancing for the latter.
* They also calculate monthly alphas for each strategy based on five different factor models, ranging from the Capital Asset Pricing Model (CAPM) to the Fama-French five-factor model plus momentum.

<mark style="color:green;">**Results**</mark>

a. Sharpe ratios:

* Equal-weighted portfolios based on GPT-4 predictions achieve a Sharpe ratio of 3.36, substantially higher than ANN-based (2.54) and logistic regression-based (2.05) portfolios.
* For value-weighted portfolios, ANN performs better (Sharpe ratio of 1.79) than GPT-4 (1.47), with both outperforming logistic regressions (0.81).

b. Alphas:

* Equal-weighted portfolios generate higher alphas in general.
* After controlling for five factors and momentum, equal-weighted portfolios based on GPT-4 predictions generate a monthly alpha of 84 basis points (10% annually), higher than ANN (60 basis points) and logistic regression (43 basis points) strategies.
* For value-weighted portfolios, ANN-based portfolios perform better than GPT-4, with monthly alphas of 50 basis points and 37 basis points, respectively, after controlling for five factors and momentum.

c. Cumulative returns:

* The authors plot the cumulative log returns of equal-weighted portfolios based on GPT-4 predictions from 1968 to 2021, showing that the long portfolio substantially outperforms the short portfolio.
* The long-short portfolio consistently outperforms the market portfolio, even when the market experiences negative cumulative returns.

The results demonstrate the potential value of GPT-4-based fundamental analysis in stock markets.&#x20;

The stronger performance of GPT-4 compared to ANN for equal-weighted strategies and the weaker performance for value-weighted strategies *<mark style="color:yellow;">**suggest that GPT-4 may have an advantage in uncovering value in smaller stocks**</mark>*.&#x20;

This finding is consistent with the authors' earlier results showing that GPT-4 appears to have an edge in analysing smaller and relatively more volatile companies.

Overall, this section provides compelling evidence for the practical utility of GPT-4-based financial statement analysis in generating profitable trading strategies.&#x20;

The authors' approach to portfolio formation and performance evaluation is comprehensive and well-justified, making a strong case for the potential of large language models in enhancing investment decision-making.

### <mark style="color:purple;">Example Output</mark>

The example output provided by GPT-4 in Appendix C demonstrates the model's step-by-step approach to analysing financial statements and predicting future earnings.

Let's break down the analysis:

<mark style="color:blue;">**Panel A. Trend Analysis:**</mark>

* GPT-4 identifies significant trends in the company's financial performance over the past three years.
* The model notes a consistent upward trend in sales, indicating strong market demand for the company's products or services.
* However, the cost of goods sold has also increased substantially, potentially affecting profitability.
* GPT-4 observes that the gross profit has increased, albeit at a slower pace, suggesting that the company has been able to maintain a degree of pricing power or cost efficiency.

#### <mark style="color:blue;">Panel B. Ratio Analysis:</mark>

* GPT-4 calculates and interprets key financial ratios to assess the company's performance.
* The model notes an improvement in the operating margin, which indicates better cost management or increased efficiency.
* GPT-4 also points out the increase in the asset turnover ratio, suggesting improved efficiency in utilizing assets.
* The model highlights a relative decline in sales efficiency, which could be a concern and may require further investigation.
* GPT-4 concludes that the company shows potential for improved profitability if cost management is maintained.

<mark style="color:blue;">**Panel C. Rationale:**</mark>

* GPT-4 synthesises the insights from the trend and ratio analyses to make a prediction about future earnings.
* The model expects the company to effectively manage its operating expenses, driven by the observed revenue growth trend and the improvement in operating margin.
* GPT-4 introduces some uncertainty into the prediction by acknowledging that the magnitude of EPS growth depends on the company's ability to maintain a degree of pricing power or cost efficiency.
* The model predicts a "moderate" change in EPS, showing potential for improved profitability, with a prediction certainty of 0.7.

The example output demonstrates GPT-4's ability to perform a structured analysis of financial statements, identifying key trends, calculating and interpreting relevant ratios, and synthesising the information to make a prediction about future earnings.&#x20;

The model's reasoning closely resembles the thought process of a human analyst, considering both quantitative and qualitative factors.

Interestingly, GPT-4 not only provides a binary prediction (increase or decrease) but also offers insights into the magnitude of the expected change and the level of certainty associated with its prediction. This additional information could be valuable for decision-makers in assessing the reliability and potential impact of the model's predictions.

Overall, the example output showcases GPT-4's capability to generate meaningful and nuanced insights from financial statements, highlighting its potential as a tool for enhancing financial analysis and decision-making.&#x20;

The model's step-by-step approach and clear articulation of its reasoning make the output interpretable and actionable, which are crucial factors in the practical application of AI-based financial analysis.


# The Evolution of AI Agents and Their Potential for Augmenting Human Agency

In this transcript, Maya Akim, an AI content creator and agent builder, discusses various aspects of AI, focusing on generative AI, large language models (LLMs), and AI agents.&#x20;

She shares her insights and experiences, highlighting the potential and challenges of these technologies.  This is an excellent analysis of the field, and well worthwhile watching.

{% embed url="<https://www.youtube.com/watch?v=fsIipBuM4Nc>" %}
Maya Akim, an AI content creator and agent builder
{% endembed %}

### <mark style="color:purple;">Introduction</mark>

The rapid advancement of artificial intelligence (AI) technology has led to the development of increasingly sophisticated <mark style="color:blue;">**AI agents**</mark> that are capable of autonomously performing tasks and making decisions.&#x20;

This video explores the philosophical underpinnings of agency and action, tracing the historical development of AI agents, and discussing how AI can be leveraged to augment human agency in various domains.

### <mark style="color:purple;">The Philosophy of Action and Agency</mark>

The concept of agency has been a subject of philosophical inquiry for centuries. Aristotle, in his work "Nicomachean Ethics," asserted that *<mark style="color:yellow;">**humans deliberate not about ends, but about means**</mark>*.&#x20;

This implies that *<mark style="color:yellow;">**the road between one's current state and desired goal consists of a series of actions.**</mark>*

In the 13th century, Ramon Llull developed logical operations using mechanical wheels, foreshadowing early computing.  Later, Blaise Pascal's mechanical calculator and Ada Lovelace's groundbreaking algorithm further advanced the notion of machines performing tasks.

The 20th century saw significant developments in the philosophy of action and agency. Alan Turing's seminal paper, "Computing Machinery and Intelligence," posed the question, "Can machines think?" and introduced the famous Turing test.&#x20;

The 1956 Dartmouth Conference coined the term "artificial intelligence," with its attendees optimistically proposing a 10-month study to make machines capable of using language, forming abstractions, and solving problems

### <mark style="color:purple;">The Evolution of AI Agents</mark>

Early AI systems, such as MYCIN[^1] in the 1970s, relied on *<mark style="color:yellow;">**symbolic AI and knowledge-based systems**</mark>*.&#x20;

These systems employed <mark style="color:blue;">rule-based reasoning</mark>, where a *<mark style="color:yellow;">**set of if-then statements**</mark>* guided the AI's decision-making process.  However, this approach proved limited due to the inherent uncertainty, ignorance, and complexity of real-world problems.

The 1980s marked a paradigm shift in AI, with the rise of *<mark style="color:yellow;">**probabilistic and statistical methods**</mark>*, deep learning, and reinforcement learning.&#x20;

This new approach enabled AI agents to learn from their environment, adapting their behaviour based on rewards and punishments.  The development of OpenAI's gym environments in 2016 further accelerated the training of AI agents in simulated environments.

Modern AI agents possess several key characteristics, including *<mark style="color:yellow;">**autonomy, memory, reactivity, proactivity, and social ability**</mark>*.&#x20;

They can perceive their environment, make decisions based on their goals and knowledge, and interact with other agents and humans to accomplish tasks.

### <mark style="color:purple;">Augmenting Human Agency with AI</mark>

AI agents have the potential to significantly augment human agency by performing tasks at a scale and speed that humans cannot match.&#x20;

For example, an AI agent can analyse vast amounts of data from various sources, such as social media, news articles, and academic papers, and provide summarised insights in a matter of minutes – a task that would take a human hours or even days to complete.

AI agents can also assist with decision-making by considering a wide range of possibilities and scenarios.&#x20;

In complex domains such as finance, healthcare, and strategic planning, AI agents can help humans make more informed decisions by analysing data, identifying patterns, and predicting outcomes.

However, it is crucial to recognize that *<mark style="color:yellow;">**AI agents are not a replacement for human judgment and oversight.**</mark>*  As AI technology continues to advance, it is essential to develop robust frameworks for human-AI collaboration, ensuring that AI agents are aligned with human values and goals.

### <mark style="color:purple;">Conclusion</mark>

The development of AI agents has its roots in centuries of philosophical inquiry into the nature of agency and action.&#x20;

As AI technology progresses, AI agents are becoming increasingly capable of autonomously performing tasks and augmenting human agency.&#x20;

By leveraging the unique capabilities of AI agents, humans can make more informed decisions, tackle complex problems, and achieve goals more efficiently.&#x20;

However, it is critical to ensure that the development and deployment of AI agents are guided by ethical principles and human oversight to maximise their benefits while mitigating potential risks.

[^1]: MYCIN was an early [backward chaining](https://en.wikipedia.org/wiki/Backward_chaining) [expert system](https://en.wikipedia.org/wiki/Expert_system) that used [artificial intelligence](https://en.wikipedia.org/wiki/Artificial_intelligence) to identify bacteria causing severe infections, such as [bacteremia](https://en.wikipedia.org/wiki/Bacteremia) and [meningitis](https://en.wikipedia.org/wiki/Meningitis), and to recommend [antibiotics](https://en.wikipedia.org/wiki/Antibiotic), with the dosage adjusted for patient's body weight — the name derived from the antibiotics themselves, as many antibiotics have the suffix "-mycin". The Mycin system was also used for the diagnosis of blood clotting diseases. MYCIN was developed over five or six years in the early 1970s at [Stanford University](https://en.wikipedia.org/wiki/Stanford_University). It was written in [Lisp](https://en.wikipedia.org/wiki/Lisp_programming_language) as the doctoral dissertation of [Edward Shortliffe](https://en.wikipedia.org/wiki/Edward_Shortliffe) under the direction of Bruce G. Buchanan, [Stanley N. Cohen](https://en.wikipedia.org/wiki/Stanley_N._Cohen) and others.


# Better Call Saul - SaulLM-7B - a legal large language model

This <mark style="color:blue;">**March 2024**</mark> paper introduces SaulLM-7B, a large language model (LLM) designed for the legal domain.&#x20;

The authors argue that while LLMs have made significant advancements in various fields, the legal domain has yet to fully benefit from this technology.&#x20;

Legal professionals are faced with an increasing volume of complex documents, and there is a growing need for a dedicated LLM to help navigate and interpret legal material.

{% embed url="<https://arxiv.org/abs/2403.03883>" %}
SaulLM-7B
{% endembed %}

### <mark style="color:purple;">Main Contributions</mark>

#### <mark style="color:green;">**A family of legal LLMs**</mark>

The authors introduce SaulLM-7B, a 7-billion-parameter language model trained on a large and diverse legal dataset. They also release SaulLM-7B-Instruct, an instruction-tuned variant that outperforms existing models on various legal tasks.

#### <mark style="color:green;">An improved evaluation protocol for legal LLMs</mark>

The authors introduce LegalBench-Instruct, an iteration of LegalBench designed to better assess the legal proficiency of language models. They also include legal tasks from the MMLU benchmark in their evaluation protocol.

### <mark style="color:purple;">The methodology for creating SaulLM-7B involves a two-step process</mark>

#### <mark style="color:green;">Enhancing Mistral's Legal Capabilities</mark>

The authors choose Mistral 7B, a high-performing open-source model, as the backbone for SaulLM-7B.&#x20;

They curate a *<mark style="color:yellow;">**high-quality legal dataset containing 30 billion tokens**</mark>* and perform continued pretraining to enhance the model's performance on legal tasks.

#### <mark style="color:green;">Improving Legal Instruction Following</mark>

To support user requests and conversational interaction, the authors fine-tune SaulLM-7B using both generic and legal instructions.&#x20;

The generic instructions help improve the model's understanding and following of commands, while the legal instructions cover tasks such as legal question answering and summarisation.

The authors note that many common LLMs include an additional step of aligning the model with human preferences. However, their early experiments did not show any meaningful improvement in performance, so they opted not to pursue this avenue for the present paper.

### <mark style="color:purple;">Data</mark>

In the "Data" section of the paper, the authors describe their data collection and cleaning processes for both the legal pretraining corpora and the instruction fine-tuning datasets.

#### <mark style="color:green;">Legal Pretraining Corpora</mark>

The authors collected legal texts from various English-speaking jurisdictions, including the U.S., Europe, and Australia, to *<mark style="color:yellow;">**capture the diversity of legal systems**</mark>*.&#x20;

They combined previously available datasets, such as subsets from The Pile and MultiLegal Pile, with data scraped from publicly available sources on the Web.&#x20;

The sources included FreeLaw, EDGAR, English EuroParl, GovInfo, Law Stack Exchange, Open Australian Legal Corpus, EULegislation, UKLegislation, Court Transcripts, and UPSTO. &#x20;

These sources contained noise and duplicated documents, which were filtered and deduplicated, resulting in a <mark style="color:blue;">**30 billion token dataset**</mark>.

To reduce the risk of catastrophic forgetting during continued pretraining, the authors incorporated "general" data from Wikipedia, StackExchange, and GitHub, comprising roughly 2% of the final training mix.&#x20;

Additionally, they included conversational data from the Super Natural Instruction and FLAN collection during pretraining, inspired by recent advances in neural machine translation.

The authors employed various data cleaning techniques to address issues in the collected data, such as text normalisation, rule-based filtering, and perplexity filtering.&#x20;

They also removed duplicates and near-duplicates using a deduplication tool, resulting in a high-quality 30B token dataset.

#### <mark style="color:green;">Instruction Fine-tuning Mixes</mark>

The authors emphasise the *<mark style="color:yellow;">**importance of instruction fine-tuning for optimal performance across different tasks.**</mark>* They used a mix of general and legal instructions to train the model, with a focus on legal expertise.

For general instructions, they gather data from four primary sources:

<mark style="color:blue;">**SlimOrca:**</mark> A subset of the [FLAN collection](#user-content-fn-1)[^1] comprising generic instructions for various tasks.

<mark style="color:blue;">**Meta Math Question Answering Instructions:**</mark> A dataset designed for mathematical inquiry, facilitating research in math-based natural language processing.

<mark style="color:blue;">**General Conversations from UltraChat:**</mark> A GPT-derived dataset capturing diverse conversational contexts to enhance natural language understanding and generation.

### <mark style="color:purple;">Evaluation</mark>

They employed three main benchmarks:

<mark style="color:green;">**Perplexity Measurement**</mark>

The authors evaluate the adaptability of the model to various legal documents by measuring [<mark style="color:blue;">**perplexity**</mark> ](mailto:undefined)on benchmark datasets from four distinct legal domains: <mark style="color:yellow;">contracts</mark>, <mark style="color:yellow;">judicial decisions</mark>, <mark style="color:yellow;">opinion text</mark>, and <mark style="color:yellow;">legislation</mark>.&#x20;

They ensure the datasets are up-to-date and sourced after the collection cut-off date from LLM data to avoid data leakage.

<mark style="color:green;">**LegalBench-Instruct**</mark>

During their investigations, the authors found limitations in the original prompts of LegalBench.&#x20;

The complex nature of the prompts, combined with the challenges faced by open-source LLMs in adhering to instructions and handling formatting, led to a substantial drop in performance.&#x20;

To address this issue, they refined the prompts by removing distracting few-shot examples and concluding with a specific instruction for the model to generate tags. This refinement aimed to provide a more accurate assessment of the model's performance on legal tasks.

#### <mark style="color:green;">Massive Multitask Language Understanding (MMLU)</mark>

The authors also used the legal section of the MMLU benchmark to gain additional insights into the model's legal knowledge. They focused specifically on three legal domains: international law, professional law, and jurisprudence.

The authors used balanced accuracy as the primary metric for evaluating the model's performance on both LegalBench-Instruct and the legal tasks of MMLU.  Balanced accuracy is chosen to better handle imbalanced classification tasks present in both benchmarks.

By using diverse benchmarks and refining the evaluation process, they aimed to provide a comprehensive understanding of the model's strengths and limitations in the legal domain. The use of up-to-date datasets and the focus on specific legal domains further enhance the validity and relevance of their findings.

### <mark style="color:purple;">Experimental Setting</mark>

#### <mark style="color:green;">Baselines</mark>

The authors compare the SaulLM-7B family to other 7B and 13B open-source models, including instruction-tuned and DPO-finetuned variants of Mistral-7B, zephyr-7b-beta, and the Llama2 family (Llama2-7b-Chat and Llama2-13b-Chat).

#### <mark style="color:green;">Implementation Details</mark>

The codebase is built using open-source frameworks like PyTorch, DeepSpeed, and Flash Attention.

The models are available on the Huggingface hub.  Continuous pretraining uses 256 MI250 AMD GPUs, while instruction fine-tuning is distributed across 16 MI250 GPUs. Evaluation is conducted on a single MI250 GPU.

#### <mark style="color:green;">Results on Legal-MMLU</mark>

SaulLM-7B-Instruct consistently outperforms non-legal instruction-tuned models on the three legal tasks of MMLU, confirming its strong performance in the legal domain.

#### <mark style="color:green;">Perplexity Analysis</mark>

SaulLM-7B *<mark style="color:yellow;">**consistently outperforms Mistral-7B across all legal document categories**</mark>*, exhibiting lower average perplexity scores with reduced variance.&#x20;

Llama2-7B demonstrates lower perplexity specifically in legislation documents, suggesting a potentially higher proportion of legislative text in its training corpora.

Overall, the experimental setting and results demonstrate the effectiveness of SaulLM-7B and SaulLM-7B-Instruct in the legal domain, establishing them as strong foundations for building models tailored to legal workflows.&#x20;

The authors provide a comprehensive analysis of their models' performance compared to other state-of-the-art open-source models, highlighting the benefits of legal-specific pretraining and instruction fine-tuning.

### <mark style="color:purple;">Conclusion</mark>

In conclusion, SaulLM-7B and SaulLM-7B-Instruct represent an advancement in the application of large language models to the legal domain.&#x20;

By leveraging extensive pretraining on legal corpora and incorporating legal-specific instruction fine-tuning, these models demonstrate superior performance on legal benchmarks such as LegalBench-Instruct and Legal-MMLU compared to generic open-source models.&#x20;

SaulLM-7B and SaulLM-7B-Instruct serve as strong foundations for building models tailored to legal workflows, paving the way for further innovation and adoption of AI in the legal field.&#x20;

[^1]: FLAN (Fine-tuned LAnguage Net) collection refers to a set of models that have been fine-tuned for a variety of tasks using a technique known as instruction tuning. This approach involves training language models to follow instructions embedded within the input data, enhancing their ability to perform specific tasks as directed by those instructions.&#x20;


# MentaLLaMA: Interpretable Mental Health Analysis on Social Media with Large Language Models

Kailai Yang, Tianlin Zhang, Ziyan Kuang, Qianqian Xie, Jimin Huang, Sophia Ananiadou

This <mark style="color:blue;">February 2024</mark> paper addresses the problem of automatically analysing mental health conditions from social media posts in an interpretable manner using large language models (LLMs).

The introduction provides an overview of the current state of mental health analysis on social media and the limitations of existing methods. &#x20;

Traditional discriminative methods, such as pre-trained language models (PLMs), achieve state-of-the-art performance in mental health-related text classification tasks.  However, these methods often struggle with poor generalisation to unseen tasks, lack robustness in multi-task scenarios, and provide predictions with low interpretability.

{% embed url="<https://arxiv.org/abs/2309.13567>" %}
MentaLLaMA: Interpretable Mental Health Analysis on Social Media with Large Language Models
{% endembed %}

To overcome these limitations, the authors explore the use of recent LLMs, such as ChatGPT and GPT-4, for interpretable mental health analysis on social media.&#x20;

These models have demonstrated superior generalisation capabilities and can provide detailed explanations for their decisions.&#x20;

However, closed-source *<mark style="color:yellow;">**LLMs like ChatGPT still fail to achieve comparable performance to state-of-the-art supervised methods in zero-shot or few-shot learning settings,**</mark>* and the low precision significantly affects the quality of the generated explanations.

The authors identify two key challenges in improving LLMs for interpretable mental health analysis through fine-tuning:

1. <mark style="color:yellow;">Lack of high-quality supervised training data</mark> that provides detailed and reliable explanations for detection results.
2. <mark style="color:yellow;">Absence of open-source LLMs specifically designed</mark> for interpretable mental health analysis.

To address these challenges, the authors formally model interpretable mental health analysis as a text-generation task, aiming to detect evidence of mental health conditions in social media posts and generate explanations for the predictions.&#x20;

They build the first multi-task and multi-source Interpretable Mental Health Instruction (IMHI) <mark style="color:yellow;">dataset with 105K data samples to support LLM instruction tuning and evaluation</mark>.&#x20;

The dataset is created by collecting raw data from various sources, using ChatGPT to generate explanations, and transforming the data into instruction-based query-answer pairs.

### <mark style="color:purple;">Raw Data Collection</mark>&#x20;

The authors collect raw data from 10 existing mental health analysis datasets spanning multiple social media sources, including Reddit, Twitter, and SMS texts.&#x20;

These datasets come with high-quality annotations, which are crucial for explanation generation and AI-generated content evaluation.&#x20;

The tasks covered in these datasets include:

<mark style="color:blue;">**Binary mental health detection:**</mark> Identifying symptoms of a single mental health condition, with binary labels. Datasets used: Depression\_Reddit (DR), CLPsych15 (CLP), Dreaded (stress detection), and a loneliness symptom detection dataset.

<mark style="color:blue;">**Multi-class mental health detection:**</mark> Identifying symptoms of one mental health condition from a given list of multiple conditions, modelled as a multi-class single-label classification task. Datasets used: T-SID and SWMH, covering depression, PTSD, anxiety, etc.

<mark style="color:blue;">**Mental health cause/factor detection:**</mark> Assigning a label to a post showing a mental health condition to identify a possible cause/factor from a given list. Datasets used: SAD (stress cause detection) and CAMS (depression/suicide cause detection).

<mark style="color:blue;">**Mental risk/wellness factors detection**</mark><mark style="color:blue;">:</mark> Identifying psychological risk/wellness factors from social media posts, modeled as a classification task. Datasets used: IRF (interpersonal risk factors) and MultiWD (mental wellness dimensions).

### <mark style="color:purple;">Explanation Generation with ChatGPT</mark>&#x20;

Due to the lack of open-source data providing detailed explanations for the annotations, *<mark style="color:yellow;">**the authors leverage ChatGPT to generate explanations.**</mark>*&#x20;

They ask domain experts to write 1 task-specific instruction and 35 explanation examples for each task, resulting in a gold explanation set G with 350 samples.&#x20;

The explanations follow a template: "\[label] Reasoning: \[explanation]".

For each dataset, they randomly sample 2 explanations per class from G as few-shot examples and include supervised annotations from the raw datasets to construct prompts for ChatGPT to generate explanations.

<mark style="color:green;">Explanation Evaluation</mark>&#x20;

The authors perform automatic and human evaluations to ensure the quality of the ChatGPT-generated explanations.

<mark style="color:green;">Automatic Evaluation</mark>&#x20;

Three criteria are used:

1. Correctness: Explanations should make correct label predictions.
2. Consistency: Explanations should provide consistent analyses with the predicted labels.
3. Quality: Explanations should provide supportive evidence with high reliability and professionality.

For correctness, they compare dataset annotations with ChatGPT responses.&#x20;

For consistency, they train classifiers using the explanation-label pairs and evaluate them on test splits and the gold explanation set G.&#x20;

For quality, they compare the generated explanations with zero-shot, few-shot, and expert-written prompts using BART-score.

#### <mark style="color:green;">Human Evaluation</mark>&#x20;

200 randomly selected explanations are assessed by domain experts on 4 aspects: consistency, reliability, professionality, and overall effectiveness, using a 0-3 rating scale.

#### <mark style="color:green;">Instruction Construction</mark>&#x20;

The IMHI dataset is constructed using the posts from raw datasets and the evaluated ChatGPT-generated explanations.&#x20;

Simplified instructions are used to adapt to less powerful LLMs. The *<mark style="color:yellow;">**training split consists of 72,095 samples, while the validation split has 14,346 samples.**</mark>*&#x20;

An IMHI-completion dataset is also created using a different template for baseline models with poor instruction-following ability.

### <mark style="color:purple;">Training Process</mark>

In this section, the authors describe the training process for their MentaLLaMA models using the IMHI dataset and the LLaMA2 models as the base.

#### <mark style="color:green;">MentaLLaMA-7B</mark>

* They finetune the LLaMA2-7B model on the IMHI training set for 10 epochs.
* The best model is selected based on the validation results on the IMHI validation set.
* Training hyperparameters:
  * Batch size: 32
  * Gradient accumulation steps: 8 (leading to an effective batch size of 256)
  * Optimizer: AdamW
  * Max learning rate: 1e-5
  * Warm-up ratio: 3%
  * Max model input length: 2048
* They use Flash-Attention to speed up the training process.

#### <mark style="color:green;">MentaLLaMA-chat-7B and MentaLLaMA-chat-13B</mark>

* These models are trained on <mark style="color:blue;">**LLaMA2-chat-7B**</mark> and <mark style="color:blue;">**LLaMA2-chat-13B**</mark><mark style="color:blue;">,</mark> respectively.
* LLaMA2-chat models are optimised with instruction tuning and are the first open-source LLMs tuned with reinforcement learning from human feedback (RLHF).
* The training process and experimental settings are the same as for MentaLLaMA-7B.

#### <mark style="color:green;">LLaMA2-7B on IMHI-completion dataset</mark>

* To enable fair comparisons with baseline models that are fine-tuned in a completion-based manner, they train another LLaMA2-7B model on the IMHI-completion dataset.

All models are trained on 4 Nvidia Tesla A100 GPUs, each with 80GB of memory.

### <mark style="color:purple;">Results</mark>

In this section, the authors present the experimental results and analysis of their proposed MentaLLaMA models in comparison to various baseline models on the IMHI test set.

The authors select the following baseline models for comparison:

1. <mark style="color:blue;">**Discriminative methods:**</mark> Classification models that finetune PLMs like BERT and RoBERTa, including SOTA methods MentalBERT and MentalRoBERTa.
2. <mark style="color:blue;">**Zero-shot/few-shot methods:**</mark> Open-source LLMs LLaMA2-7B and LLaMA2-13B for zero-shot prompting, and closed-source LLMs ChatGPT and GPT-4 for zero-shot and few-shot prompting.
3. <mark style="color:blue;">**Completion-based fine-tuning methods:**</mark> SOTA generative PLMs BART-large and T5-large finetuned on the IMHI-completion dataset, along with a LLaMA-7B model for fair comparison.

### <mark style="color:purple;">IMHI Test Results</mark>

#### <mark style="color:green;">Correctness</mark>

* MentalBERT and MentalRoBERTa achieve SOTA performance on 8 out of 10 test sets among discriminative methods.
* ChatGPT significantly outperforms LLaMA2 models in zero-shot settings, and few-shot learning further improves its performance.
* <mark style="color:yellow;">Fine-tuning methods show significant improvement over LLaMA2 zero-shot results.</mark>
* MentaLLaMA-7B outperforms completion-based LLaMA2-7B on 8 out of 10 test sets, *<mark style="color:yellow;">**showing the efficiency of domain-specific instruction tuning.**</mark>*
* MentaLLaMA-chat-13B surpasses or closely matches MentalRoBERTa in 7 out of 10 test sets.

#### <mark style="color:green;">Quality</mark>

* In completion-based methods, LLaMA2-7B outperforms its zero-shot counterpart, and BART-large is recommended for building a completion-based interpretable mental health analysis model.
* In instruction tuning methods, MentaLLaMA greatly outperforms zero-shot results on LLaMA2-7B, and MentaLLaMA-chat models further improve the quality of explanations.
* MentaLLaMA models achieve comparable performance to ChatGPT and GPT-4 on the expert-written gold set with much smaller model sizes.

#### <mark style="color:green;">**Generalisability**</mark>

* MentaLLaMA models significantly outperform LLaMA2-13B and ChatGPT in zero-shot settings on unseen tasks.
* MentaLLaMA-chat models generate higher quality explanations compared to T5 and BART on unseen tasks, especially in mental health conditions/cause detection and high-level mental health factors.
* MentaLLaMA-chat-13B further improves the explanation quality compared to MentaLLaMA-chat-7B, showing the benefit of model size expansion.

#### <mark style="color:green;">**Human Evaluation**</mark>

* MentaLLaMA-chat-13B achieves high scores on consistency and reliability, comparable to ChatGPT.
* However, MentalLLaMA underperforms ChatGPT in professionality, indicating a lack of domain-specific knowledge. Continual pre-training on high-quality mental health-related data is suggested as a solution.

### <mark style="color:purple;">Conclusion</mark>

Overall, the results demonstrate the effectiveness of the MentaLLaMA models in achieving high correctness, generating quality explanations, and exhibiting strong generalisability to unseen tasks in the domain of interpretable mental health analysis on social media.

In this paper, the authors introduce the novel task of interpretable mental health analysis and present the first multi-task and multi-source dataset, IMHI, which contains 105K data samples for instruction tuning.

They used ChatGPT to generate the training data and perform rigorous automatic and human evaluations to ensure its reliability.&#x20;

Building upon the IMHI dataset, the authors propose MentaLLaMA, the first open-source large language model series designed for interpretable mental health analysis with instruction-following capabilities.

Evaluations on the IMHI benchmark demonstrate that MentaLLaMA achieves performance comparable to state-of-the-art discriminative methods in terms of correctness and generates explanations on par with human-level quality.  Additionally, MentaLLaMA exhibits strong generalisability to unseen tasks.

However, the authors acknowledge that *<mark style="color:yellow;">**MentaLLaMA still lacks domain-specific knowledge compared to powerful models like ChatGPT.**</mark>*&#x20;

Future work will explore *<mark style="color:yellow;">**continual pre-training of MentaLLaMA on large-scale, high-quality mental health-related data to enhance the professionality of its explanations.**</mark>*&#x20;


# Anomaly detection in logging data

This 2024 paper introduces <mark style="color:blue;">**LogFiT**</mark>, a novel log anomaly detection model that leverages the power of pretrained language models (LMs) like RoBERTa and Longformer to identify anomalies in system logs.&#x20;

The key idea behind LogFiT is to fine-tune these LMs on normal log data, enabling them to learn the linguistic and sequential patterns of normal logs. By doing so, the model can then detect anomalies when presented with new log data that deviates from these learned patterns.

{% file src="/files/v9aKb5iT2nH9DYmRp1DV" %}

### <mark style="color:purple;">Issues in Log Data Collection and Processing</mark>

<mark style="color:green;">**Variability in Log Content**</mark>

* Traditional log anomaly detection systems struggle with the inherent variability in log entries. Logs are dynamic, with their formats and contents changing as the underlying systems are updated or configured differently. This variability can render static log parsing methods ineffective, as they fail to adapt to new log patterns and structures.

<mark style="color:green;">**Dependence on Log Templates**</mark>

* Many existing systems rely on predefined log templates to parse and interpret log data. This method is inflexible, as it requires logs to fit these templates exactly. When logs deviate from expected formats—common due to system upgrades or changes—the template-based methods fail to recognize important log entries, leading to missed anomalies.

<mark style="color:green;">**Need for Labeled Data**</mark>

* Supervised learning approaches, which are prevalent in many machine learning applications, require extensively labeled datasets to train models effectively. In the context of log analysis, obtaining a sufficiently large and accurately labeled dataset is often costly and time-consuming. This requirement limits the practical deployment of sophisticated machine learning models in operational environments where labeled data is scarce.

<mark style="color:green;">**Self-Supervised Learning Limitations**</mark>

* The paper discusses two types of self-supervised learning models: forecasting-based and reconstruction-based. Both types attempt to learn the normal patterns of log entries to detect anomalies. However, these models traditionally require substantial modifications when log data characteristics change, which is a common occurrence in dynamic IT environments.

### <mark style="color:purple;">Proposed Solution - LogFiT Model</mark>

The LogFiT model proposed in the paper addresses these issues by leveraging a pretrained BERT-based language model fine-tuned for understanding the linguistic patterns of normal log data.&#x20;

Key features and benefits of the LogFiT model include:

* <mark style="color:blue;">**Robustness to Changes in Log Content**</mark><mark style="color:blue;">:</mark> By using a language model trained on a broad corpus of text, LogFiT can adapt to changes in log syntax and vocabulary without requiring retraining or extensive manual adjustments.
* <mark style="color:blue;">**Self-Supervised Training on Normal Data**</mark><mark style="color:blue;">:</mark> LogFiT is trained using masked token prediction on normal log data, eliminating the need for labeled anomaly data. This training approach makes it suitable for environments where anomaly labels are unavailable.
* <mark style="color:blue;">**High Performance Across Diverse Datasets**</mark><mark style="color:blue;">:</mark> The model demonstrates superior F1 scores compared to traditional methods like DeepLog and LogBERT, particularly when log data variability is introduced during evaluation. This indicates that LogFiT can maintain high accuracy even as the characteristics of log data evolve.

#### <mark style="color:green;">Self-supervised training</mark>

LogFiT is trained using only normal log data in a self-supervised manner. It does not require any labeled data, making it more practical for real-world scenarios where labeled anomalies are scarce.

#### <mark style="color:green;">Masked sentence prediction</mark>

The model is trained using a novel masked sentence prediction objective. It randomly masks a variable ratio of sentences and tokens within each log paragraph and learns to predict the masked tokens. This approach helps the model learn the contextual relationships between tokens and sentences, enabling it to understand the language rules of normal logs.

<figure><img src="/files/zK2u1FudeOPQIOhzr3JI" alt=""><figcaption><p>HDFS log sentences converted to log templates</p></figcaption></figure>

#### <mark style="color:green;">Pretrained LMs</mark>

LogFiT leverages pretrained LMs like RoBERTa and Longformer, which have been shown to capture both syntactic and semantic information. The model selects between RoBERTa and Longformer based on the length of log sequences, with Longformer being used for sequences exceeding 512 tokens.

#### <mark style="color:green;">Fine-tuning</mark>

The pretrained LMs are fine-tuned on normal log data using techniques like gradual unfreezing and super-convergence. This fine-tuning process adapts the LMs to the specific domain of system logs, improving their ability to detect anomalies.

#### <mark style="color:green;">Anomaly detection</mark>

During inference, LogFiT uses the fine-tuned model's top-k prediction accuracy as an anomaly score. If the model's accuracy in predicting masked tokens falls below a certain threshold, the log paragraph is considered an anomaly.

### <mark style="color:purple;">Process of log anomaly detection</mark>

The paper elaborates on the <mark style="color:yellow;">process of log anomaly detection</mark> which is broken down into several detailed steps.   Each step addresses specific challenges related to handling and processing log data for anomaly detection purposes.&#x20;

#### <mark style="color:blue;">1. Log Data Pre-processing</mark>

This initial stage involves cleaning and standardising raw system logs.&#x20;

Since logs are generated by various system processes and applications, they often contain a lot of noise such as redundant data, irrelevant information, and inconsistencies.&#x20;

The goal of this step is to format the data uniformly to ensure the reliability and accuracy of analysis in subsequent stages. Pre-processing might include filtering out unnecessary information, correcting or standardizing time formats, or resolving ambiguities in log entries.

#### <mark style="color:blue;">2. Vectorization</mark>

After the logs are cleaned and standardised, they are <mark style="color:yellow;">converted into numerical representations known as vectors</mark>.&#x20;

This process is important because machine learning models, particularly those based on deep learning, require numerical input to perform computations.&#x20;

The method of vectorisation can vary: basic techniques might use one-hot encoding or frequency-based methods, whereas more advanced approaches might employ semantic vectorization techniques.&#x20;

Semantic vectorization involves embedding words or phrases from the logs into vectors that capture not just the occurrence of terms but also their meanings based on the context in which they appear.

#### <mark style="color:blue;">3. Model Development</mark>

In this phase, the data that has been transformed into vectors is used to train a machine learning model.&#x20;

The choice of model architecture, training objectives, and evaluation metrics are critical decisions made based on the specific characteristics of the data and the requirements of the anomaly detection task.&#x20;

Models such as Long Short-Term Memory (LSTM) networks or Transformer-based architectures like BERT are commonly used because they are effective at capturing sequential dependencies and complex patterns in data. The training process involves adjusting the model parameters to minimize prediction errors, typically measured by a loss function.

#### <mark style="color:blue;">4. Model Operationalisation</mark>

The final step involves deploying the trained model into a production environment where it can start analysing new log data to detect anomalies.&#x20;

This stage requires ensuring the model integrates seamlessly with the existing IT infrastructure and operates efficiently under operational loads. Continuous monitoring and maintenance are also crucial to ensure the model adapts to changes in data patterns or system updates.

#### <mark style="color:green;">Semantic Vectors and Their Relevance</mark>

Semantic vectors play a role in transforming log data into a format suitable for deep learning models.&#x20;

Unlike traditional numerical vectors, semantic vectors encapsulate the meanings of words or phrases within their dimensional attributes. This semantic embedding allows models to understand and interpret the context and nuances of log entries, which is essential for accurately identifying anomalies that may indicate operational issues or security threats.

#### <mark style="color:green;">Integration with Vector Databases</mark>

Vector databases can be integrated into this process as they are designed to efficiently store and manage vector data. In the context of log anomaly detection, vector databases can be used to store semantic vectors of log entries and enable fast retrieval and comparison of log patterns.&#x20;

By leveraging similarity search capabilities of vector databases, systems can quickly identify log entries that deviate from normal patterns, thereby enhancing the efficiency and responsiveness of anomaly detection systems.

By following this process and utilising advanced techniques like semantic vectorization and vector databases, organisations can significantly improve their ability to detect and respond to anomalies in log data, thereby enhancing their overall security and operational efficiency.

The architecture selection for the LogFiT model was influenced by the need to address specific challenges in log anomaly detection, particularly those related to handling large sequences of log data and capturing their complex linguistic and sequential patterns.&#x20;

### <mark style="color:purple;">Decision Process for Model Architecture</mark>

<mark style="color:blue;">**Foundation Model Selection**</mark>

The choice to use Longformer, a derivative of RoBERTa which itself builds on BERT, was strategic. Longformer was selected due to its ability to handle input sequences much longer than the 512-token limit imposed by BERT. &#x20;

This capability is critical for processing extensive log entries that contain vital sequential information for anomaly detection.

<mark style="color:blue;">**Linguistic and Sequential Pattern Learning**</mark>

The model needed to effectively learn and interpret both the linguistic structure and the sequence of events in log data. Longformer’s architecture supports this requirement by processing sequences up to 4096 tokens, allowing for comprehensive analysis of longer log entries.

<mark style="color:blue;">**Adaptation for Log Data**</mark>

The standard Longformer model was fine-tuned specifically for the domain of log data. This fine-tuning involved adjusting the model to focus on the linguistic patterns typical of log entries, which are different from the general language patterns the base model was originally trained on.

### <mark style="color:purple;">Challenges Encountered</mark>

<mark style="color:green;">**Handling High Variability in Log Data**</mark><mark style="color:green;">:</mark> Log entries can vary significantly in format, terminology, and structure, depending on the source (e.g., different software systems or hardware components). This variability makes it difficult to standardise and analyse logs effectively using a one-size-fits-all model.

<mark style="color:green;">**Integration of Semantic Understanding**</mark><mark style="color:green;">:</mark> Unlike traditional models that rely on log parsing and fixed templates, LogFiT needed to directly interpret raw log data. This required the model to not only parse the text but also understand its semantic context—an advanced requirement that necessitates deep learning capabilities.

<mark style="color:green;">**Scalability and Performance**</mark><mark style="color:green;">:</mark> Managing and processing vast amounts of log data in real-time present significant performance and scalability challenges. The model architecture needed to be efficient enough to handle large-scale data without compromising on speed or accuracy.

### <mark style="color:purple;">Insights and Interesting Outcomes</mark>

<mark style="color:green;">**Semantic Vectorization**</mark><mark style="color:green;">:</mark> LogFiT incorporates semantic vectors to represent log data, moving away from traditional log parsing methods. This approach allows the model to capture more nuanced information and adapt to changes in log data over time.

<mark style="color:green;">**Self-supervised Training Approach**</mark><mark style="color:green;">:</mark> By adopting a masked language modeling approach for training, LogFiT can effectively learn from 'normal' log data without needing labeled examples of anomalies. This method helps the model understand what typical log data should look like and identify deviations based on learned patterns.

<mark style="color:green;">**Anomaly Detection Efficacy**</mark><mark style="color:green;">:</mark> Initial experiments demonstrated that LogFiT could surpass traditional models like DeepLog and LogBERT in detecting anomalies. This improvement was particularly notable when handling logs with high variability and evolving content, which pose significant challenges for models relying on static templates.

<mark style="color:green;">**Threshold-Based Anomaly Identification**</mark><mark style="color:green;">:</mark> The use of a top-k accuracy threshold to determine anomalies provided a flexible and robust mechanism for classification. This method allowed for dynamic adjustment based on the specific sensitivity and specificity needs of different environments.

Overall, the development of LogFiT involved a thoughtful integration of advanced NLP techniques with the specific requirements of log anomaly detection. The model's ability to process long sequences and its robust training methodology contribute significantly to its effectiveness, offering a promising solution for managing the complex and dynamic nature of system logs in various IT environments.

### <mark style="color:purple;">Traditional Log Parsing Methods</mark>

Traditional log parsing methods involve transforming raw log data into a structured format by identifying and extracting predefined patterns or templates.&#x20;

This process typically includes:

<mark style="color:blue;">**Template Extraction**</mark><mark style="color:blue;">:</mark> Identifying common patterns or templates within the log data. This is often done manually or using rule-based algorithms which determine how log messages are split into fields.

<mark style="color:blue;">**Log Structuring**</mark><mark style="color:blue;">:</mark> Applying these templates to parse incoming log data into structured records. Each part of a log message is assigned to predefined fields such as timestamp, log level, message text, etc.

<mark style="color:blue;">**Fixed Format**</mark><mark style="color:blue;">:</mark> The output is a set of structured logs where each entry adheres to a fixed schema derived from the identified templates.

This approach has limitations, especially in handling log variability and evolving data formats. Since it relies on fixed templates, any deviation in log message format can lead to parsing errors or missed information.

### <mark style="color:purple;">Experimental Setup</mark>&#x20;

The selection of HDFS, BGL, and Thunderbird datasets for training and evaluating the LogFiT model is strategic:

<mark style="color:blue;">**Benchmarking**</mark><mark style="color:blue;">:</mark> These datasets are common benchmarks in the field, allowing for direct comparison with other models, particularly the baseline models like DeepLog and LogBERT. This ensures that the improvements or shortcomings of LogFiT are evident and credible within the context of existing technologies.

<mark style="color:blue;">**Real-World Relevance**</mark><mark style="color:blue;">:</mark> Each dataset represents real system logs from significant and varied computing environments (e.g., Hadoop distributed file systems, supercomputers). This choice underlines the model’s applicability to diverse operational settings.

<mark style="color:blue;">**Challenge and Complexity**</mark><mark style="color:blue;">:</mark> These datasets include both normal and anomalous log entries, with anomalies clearly labeled. This provides a robust challenge to the LogFiT model to distinguish between normal operations and potential issues, which is crucial for practical deployment in system monitoring.

### <mark style="color:purple;">Training and Evaluation Methodology</mark>

The structured approach to model training and evaluation offers insights into the rigorous methodological standards followed:

1. <mark style="color:blue;">**Self-supervised Learning**</mark><mark style="color:blue;">:</mark> By training on normal log data only, the LogFiT model leverages self-supervised learning to understand 'normal' behavioral patterns without the need for anomaly labels during training. This is particularly useful in real-world scenarios where anomalies are rare or not previously known.
2. <mark style="color:blue;">**Semantic Vectorization**</mark><mark style="color:blue;">:</mark> Directly processing log data into semantic vectors without relying on intermediate log templates allows the model to capture nuanced information and adapt to changes in log formats over time. This method shows an advanced understanding of the dynamic nature of log data.
3. <mark style="color:blue;">**Hyperparameter Tuning and Evaluation Sets**</mark><mark style="color:blue;">:</mark> The use of separate tuning and evaluation sets, and the avoidance of random shuffling to maintain the sequential integrity of log data, emphasizes the importance of evaluating the model under conditions that closely simulate actual operational environments.

### <mark style="color:purple;">Insights and Outcomes</mark>

The experimental setup and the subsequent results demonstrate that the *<mark style="color:yellow;">LogFiT model can effectively handle variations in log data syntax and semantics</mark>*, improving upon the limitations of traditional log parsing methods which rely heavily on static templates.&#x20;

The use of semantic vectors and a transformer-based architecture like Longformer allows the model to process and analyse extensive log data with high accuracy, addressing both the immediate and contextual anomalies in log sequences.

### <mark style="color:purple;">Results and Analysis</mark>

The experimental results are quite revealing, highlighting the superior performance of the LogFiT model over the baselines (DeepLog and LogBERT):

<mark style="color:blue;">**High F1 Scores and Specificity**</mark>

LogFiT consistently shows higher F1 scores across all datasets compared to the baselines, indicating a better balance between precision and recall. High specificity values suggest that LogFiT effectively reduces false positives, a crucial factor in operational settings where frequent false alarms can be disruptive.

<mark style="color:green;">**F1 Score:**</mark> The F1 score is a measure used in statistics to assess the accuracy of a test. It considers both the precision and the recall of the test to compute the score.

Precision is the ratio of correctly predicted positive observations to the total predicted positives (true positives / (true positives + false positives)).&#x20;

Recall (also known as sensitivity) is the ratio of correctly predicted positive observations to all observations in the actual class (true positives / (true positives + false negatives)).&#x20;

The F1 score is the harmonic mean of precision and recall, given by:

$$F1=2×(  Precision+Recall Precision×Recall ​  )$$

<mark style="color:yellow;">A higher F1 score indicates a more robust model that balances both precision and recall</mark> effectively, which is particularly important in systems where both the completeness (recall) and exactness (precision) of the predictions are crucial.

<mark style="color:green;">**False Positive**</mark><mark style="color:green;">:</mark> A false positive occurs when a test wrongly predicts the positive class. For example, in anomaly detection, a false positive would be an instance where the model incorrectly identifies a normal activity as an anomaly. *<mark style="color:yellow;">High rates of false positives can lead to unnecessary alarms and actions</mark>*, which can be costly and disruptive in operational environments.

<mark style="color:blue;">**Robustness to Log Data Variability**</mark>

The use of semantic vectors allows LogFiT to adapt to changes in log data over time, which is a significant improvement over traditional methods that rely on fixed log templates. This adaptability is crucial for long-term application as log formats and content can evolve.

#### Why Semantic Vectors Improve Robustness to Log Data Variability

Semantic vectors are numerical representations of text that capture the meanings of words or phrases within their dimensional attributes. These vectors go beyond simple keyword matching by understanding the context and the semantic relationships within the text, thanks to techniques like word embeddings or transformer models.

**Benefits of Semantic Vectors**:

1. **Contextual Awareness**: They capture not just the presence of words but their contextual usage within the logs. This helps in understanding the semantic similarity between different log entries even if they don't share exact words.
2. **Adaptability to Changes**: Log formats and content can evolve over time due to updates in software or changes in system configuration. Semantic vectors, because they understand context, can adapt to new terms or variations in phrasing without requiring a complete redefinition of the template or model.
3. **Generalisation**: They can generalize from the training data to unseen instances better than traditional methods, which may rely on exact match templates that fail when unexpected variations appear.

<mark style="color:blue;">**Centroid Distance Minimisation**</mark>

The experiments with centroid distance minimisation show that this method does not significantly contribute to distinguishing between normal and anomalous logs. This finding suggests that the core strength of LogFiT lies in its ability to model the normal behavior accurately without relying on the spatial relationships of log embeddings.

#### Explaining Centroid Distance Minimisation

**Centroid Distance Minimisation** is a technique often used in clustering and anomaly detection models. The centroid is the mean vector of all the points (or log entries, in this case) in a cluster, representing a sort of 'average' or typical example of the points within that cluster.

**Usage in Anomaly Detection**:

* **Centroid Calculation**: During training, the centroid of the 'normal' log entries is computed. This centroid represents the typical behavior of the system under normal conditions.
* **Distance Measurement**: During inference, the distance of a new log entry from this centroid is measured. If the distance is above a predefined threshold, the log entry is classified as anomalous, implying it is significantly different from typical behavior.
* **Minimisation Objective**: The model may also include an objective during training to minimize this distance for normal log entries, ensuring that they cluster tightly around the centroid.

The method aims to make the model sensitive to deviations from normal behavior, assuming that normal behavior can be somewhat consistently defined in terms of log entries.&#x20;

However, as noted, centroid distance minimisation did not prove effective in improving anomaly detection, likely because normal behavior itself can be diverse and not easily encapsulated by a single centroid, especially in complex systems where log data can vary significantly even under normal conditions. This highlights a limitation of using spatial relationships alone to define normalcy in dynamic environments.

### <mark style="color:purple;">Insights and Interesting Outcomes</mark>

A few key insights can be drawn from the experimental setup and results:

1. **Effectiveness of Semantic Vectorization**: The ability of LogFiT to interpret and analyse logs *<mark style="color:yellow;">without transforming them into a rigid template format</mark>* shows the power of semantic vectorization. This approach not only captures the textual content but also the context and subtle nuances, which enhances the model’s predictive accuracy.
2. **Importance of Self-Supervised Learning**: The model's training on normal data only and its subsequent performance underline the effectiveness of self-supervised learning in scenarios where anomalies are rare or not well-defined.
3. **Practical Implications**: The high specificity and F1 scores demonstrate LogFiT's potential for real-world applications, reducing the operational overhead of dealing with false positives while reliably catching anomalies.

The thoughtful experimental design and the comprehensive evaluation of the LogFiT model offer a robust proof of concept for its application in log anomaly detection. The results not only validate the model's theoretical underpinnings but also showcase its practical viability in dynamic and demanding operational environments.

### <mark style="color:purple;">Applications</mark>

Here are five applications of how LogFiT could be used on Elasticsearch logging data, detailing the type of data flow and how LogFiT could be applied:

<mark style="color:green;">**Amazon CloudWatch (Logs and Metrics)**</mark>

* **Data Flow**: Continuous streaming of log and metric data from Amazon CloudWatch into Elasticsearch.
* **LogFiT Application**: LogFiT could analyse CloudWatch logs and metrics for anomalies that indicate operational issues or potential security threats in AWS environments. By training on typical patterns of AWS resource usage and logs, LogFiT could detect deviations that signify critical incidents or inefficiencies.

<mark style="color:green;">**Apache Web Server (Logs and Metrics)**</mark>

* **Data Flow**: Real-time ingestion of Apache server logs and performance metrics into Elasticsearch.
* **LogFiT Application**: LogFiT could be used to monitor Apache logs for unusual access patterns or error rates that could indicate a web attack or system failure. By understanding normal access and error logs, LogFiT could flag anomalies for immediate action.

<mark style="color:green;">**Cisco ASA (Logs)**</mark>

* **Data Flow**: Logs from Cisco ASA firewalls are streamed into Elasticsearch for security analysis.
* **LogFiT Application**: Using LogFiT, organizations could enhance their security posture by detecting anomalies in firewall logs that might indicate breaches or unauthorized access attempts. Training on normal network traffic logs would allow LogFiT to recognize and alert on suspicious activities.

<mark style="color:green;">**Microsoft 365 Defender (Logs)**</mark>

* **Data Flow**: Security logs from Microsoft 365 Defender are collected in Elasticsearch for threat detection and analysis.
* **LogFiT Application**: LogFiT could be deployed to detect anomalies in the behavior of users and endpoints within the Microsoft 365 ecosystem, identifying potential security incidents like phishing attacks or malware infections based on deviations from baseline security logs.

<mark style="color:green;">**Docker (Logs and Metrics)**</mark>

* **Data Flow**: Docker container logs and metrics are continuously fed into Elasticsearch for monitoring container health and activity.
* **LogFiT Application**: LogFiT could be trained on normal operational logs and metrics from Docker to detect anomalies in container performance or security issues, such as containers attempting unauthorized actions. This would help in proactive container management and security enforcement.

Each of these applications involves configuring LogFiT to learn from 'normal' operational data specific to each environment, allowing it to effectively identify deviations.&#x20;

This is crucial in dynamic environments like those monitored using Elasticsearch, where log formats and schemas can frequently change and evolve. LogFiT's ability to adapt to these changes through semantic vectorization without the need for fixed log templates gives it an edge in maintaining accurate and relevant anomaly detection capabilities.


# ChatDoctor: Artificial Intelligence powered doctors

The paper presents the development and evaluation of ChatDoctor, a model fine-tuned on large language models (LLMs) specifically for the medical domain.&#x20;

The authors conducted experiments by posing medically relevant questions to ChatDoctor and assessed its performance through a blind evaluation against ChatGPT, focusing on its ability to recommend medications accurately.&#x20;

{% embed url="<https://arxiv.org/abs/2303.14070>" %}
ChatDoctor: A Medical Chat Model Fine-tuned on Llama Model using Medical Domain Knowledge
{% endembed %}

<mark style="color:yellow;">ChatDoctor demonstrated a higher accuracy (91.25%) in recommending medications</mark> based on diseases compared to ChatGPT (87.5%).

The analysis of ChatDoctor's responses to various medical inquiries <mark style="color:yellow;">revealed its potential in understanding complex medical conditions and providing appropriate recommendations</mark>.&#x20;

For instance, ChatDoctor correctly identified the need for surgical intervention in pyloric stenosis and offered medication options when surgery was not inquired about, reflecting its comprehensive understanding of medical treatment options.&#x20;

Additionally, it showcased caution by inquiring about other medications the patient might be taking before recommending treatment for myoclonus, indicating a thoughtful approach to drug interactions.

ChatDoctor also promptly recognised the urgency of carbon monoxide poisoning and advised immediate medical attention, demonstrating its ability to prioritize medical emergencies.&#x20;

However, for conditions like Wernicke-Korsakoff syndrome, ChatDoctor suggested consultation with a specialist, indicating its limitations in providing detailed advice on less common conditions.

Despite these promising results, *<mark style="color:yellow;">**the authors acknowledge significant limitations**</mark>**.***&#x20;

They emphasise that *<mark style="color:yellow;">**ChatDoctor is intended for academic research only**</mark>*, highlighting the absence of sufficient security measures, the inability to guarantee complete accuracy in medical diagnoses and recommendations, and the restriction against commercial and clinical use due to licensing constraints.

The discussion concludes with reflections on the future direction of ChatDoctor and similar models.&#x20;

The authors suggest that further improvements should focus on <mark style="color:yellow;">limiting LLMs to generate only responses with high confidence and incorporating additional safety checks</mark>, either traditional or AI-based, to mitigate the risks associated with inaccurate medical advice.&#x20;

They also note the critical need for <mark style="color:yellow;">high-quality training data to enhance model performance</mark>.&#x20;

Despite these challenges, the potential of ChatDoctor to improve medical diagnostics, *<mark style="color:green;">**reduce healthcare professionals' workload**</mark>*, and expand access to medical advice, particularly in underserved regions, is underscored as a significant contribution to healthcare and medical research.

Despite their success in generating human-like responses across general domains, these models fall short in providing accurate medical advice, diagnoses, and medication recommendations due to a lack of domain-specific training.&#x20;

To bridge this gap, the authors have collected a comprehensive dataset comprising over 700 diseases, their symptoms, necessary medical tests, and recommended medications, generating 5,000 doctor-patient conversation samples for fine-tuning LLMs.

### <mark style="color:purple;">How could this application be used?</mark>

This fine-tuning process aims to equip LLMs with the nuanced understanding required to offer informed medical advice, thereby enhancing their applicability in healthcare settings.&#x20;

The envisioned ChatDoctor model is expected to significantly improve patient care by assisting with initial diagnoses, triage, and offering medical recommendations, especially in regions with limited access to healthcare services.

The paper's main contributions are threefold:&#x20;

1. the development of a novel framework for fine-tuning LLMs in the medical domain
2. the creation of a significant dataset of doctor-patient conversations for model training
3. the demonstration of the fine-tuned model's potential for real-world clinical application.

The project represents a significant step forward in integrating advanced language models into healthcare, promising to improve the efficiency and quality of patient care by facilitating better communication between healthcare providers and patients.&#x20;

The authors have made the source codes, datasets, and model weights publicly available to encourage further research and development in this field, providing a valuable resource for the advancement of dialogue models in the medical domain.


# Navigating the Jagged Technological Frontier: Effects of AI on Knowledge Workers

Harvard Business School - 2023

The paper discusses an experiment conducted with the Boston Consulting Group to assess the impact of Artificial Intelligence (AI), particularly Large Language Models (LLMs) like GPT-4, on knowledge-intensive tasks performed by professional consultants.&#x20;

The study involved 758 consultants, who were divided into three groups: one with no AI access, one with GPT-4 access, and one with GPT-4 access plus an overview of prompt engineering.

{% embed url="<https://www.bcg.com/publications/2023/how-people-create-and-destroy-value-with-gen-ai>" %}

### <mark style="color:purple;">Key Findings</mark>

#### <mark style="color:green;">**Jagged Technological Frontier**</mark>

AI shows a varied performance across different tasks, creating a "jagged technological frontier." Some tasks are easily handled by AI, enhancing productivity and quality, while others, appearing similarly complex, are not within AI's current capabilities.

#### <mark style="color:green;">**Productivity and Quality Gains**</mark>

Consultants using AI *<mark style="color:yellow;">**completed tasks 12.2% more and 25.1% faster on average than those without AI**</mark>*.&#x20;

The quality of their work was also over 40% higher compared to the control group. These gains were observed across the skill distribution, with below-average performers showing a 43% increase and above-average performers a 17% increase in their scores.

#### <mark style="color:green;">**Detrimental Effects Outside AI Frontier**</mark>

For tasks deemed outside the AI's frontier, consultants using AI were 19% less likely to produce correct solutions, indicating that AI's utility diminishes for certain complex tasks.

#### <mark style="color:green;">**Human-AI Collaboration Models**</mark>

Two distinct patterns of AI integration emerged among consultants:

* <mark style="color:blue;">**Centaurs**</mark><mark style="color:blue;">:</mark> These consultants split tasks between themselves and AI, deciding which parts of the task to delegate to the AI and which to handle themselves.
* <mark style="color:blue;">**Cyborgs**</mark><mark style="color:blue;">:</mark> These consultants integrated AI throughout their task flow, continually interacting with the technology for a more seamless integration.

<mark style="color:green;">**Implications for Knowledge Work**</mark>

The paper underscores the significant potential of AI to transform high-skill professional environments, emphasising that AI's role in automating or augmenting tasks is not uniformly predictable across different types of work.

<mark style="color:green;">**Navigating the AI Frontier**</mark>

The study illustrates the need for professionals to develop a nuanced understanding of AI's capabilities and limitations to effectively integrate AI tools into their workflows and maximise productivity and quality gains.

### <mark style="color:purple;">Results of the experiment</mark>

#### <mark style="color:green;">**Experimental Design**</mark>

The study involved an experiment where consultants were tasked with developing innovative concepts for beverages and footwear for niche markets. These tasks were meant to simulate real-world, complex, and knowledge-intensive workflows.

#### <mark style="color:green;">**Quality and Productivity Enhancement**</mark>

When tasks were within the capabilities of AI ("inside the frontier"), consultants using AI showed significant improvements in productivity and quality. &#x20;

They completed 12.2% more tasks and did so 25.1% faster. The *<mark style="color:yellow;">**quality of their outputs, as assessed by human graders and GPT-4, was over 40% higher compared to those without AI access**</mark>*.

#### <mark style="color:green;">**Treatment Comparisons**</mark>

The study compared three groups: no AI access, GPT-4 access, and GPT-4 access with a prompt engineering overview. &#x20;

The latter two groups outperformed the no AI group, with GPT-4 plus overview showing slightly better results, suggesting that a *<mark style="color:yellow;">**structured approach to using AI can enhance its benefits**</mark>*.

#### <mark style="color:green;">**Skill Level Impact**</mark>

Both top-half and bottom-half skill performers, as identified in a preliminary assessment, benefited from AI.&#x20;

However, the bottom-half skill performers showed a more significant improvement (43%) compared to the top-half performers (17%).

#### <mark style="color:green;">**Task Completion and Speed**</mark>

Participants with AI support completed more tasks within the given time frame and did so more quickly. Specifically, the GPT + Overview group was 22.5% faster, and the GPT Only group was 27.63% faster compared to the control group.

#### <mark style="color:green;">**Diversity of Ideas**</mark>

While AI usage led to higher-quality outputs, it also resulted in less variability in the ideas generated by participants. This indicates a *<mark style="color:yellow;">**potential trade-off between quality and diversity**</mark>* when using AI for creative tasks.

#### <mark style="color:green;">**Conclusion**</mark>

The experiment underscores AI's potential to significantly enhance the performance of highly skilled workers on tasks that fall within its current capabilities. However, the use of AI needs to be nuanced, understanding that its impact varies depending on the nature of the task and the users' skill levels.

### <mark style="color:purple;">**Key Findings**</mark>

<mark style="color:green;">**Correctness in Strategic Recommendations**</mark>

The experiment's primary measure was the accuracy of strategic recommendations made by consultants.&#x20;

The control group, without AI assistance, had a correctness rate of 84.5%.&#x20;

In contrast, the groups with AI access (GPT-4 alone and GPT-4 with an overview) had lower correctness rates of 70% and 60%, respectively, *<mark style="color:yellow;">**showing a notable decline in performance when using AI for tasks outside its capability frontie**</mark>*&#x72;.

#### <mark style="color:green;">**AI's Negative Impact**</mark>

Linear regression analysis confirmed that *<mark style="color:yellow;">**AI usage negatively impacted the accuracy of solutions in this more complex, integrative task**</mark>*, with the GPT-4 + Overview group experiencing a more significant decrease in correctness.

#### <mark style="color:green;">**Efficiency**</mark>

Despite the decrease in correctness, AI usage led to faster task completion. The GPT-4 + Overview group finished tasks 30% faster, and the GPT-4 Only group was 18% faster than the control group, indicating that *<mark style="color:yellow;">**AI can increase efficiency even when it doesn't enhance task correctness**</mark>*.

#### <mark style="color:green;">**Quality of Recommendations**</mark>

Surprisingly, even when recommendations were incorrect, the quality of the advice, as evaluated by human graders, was higher in the AI-assisted groups. This suggests that while AI might lead to incorrect conclusions, the articulation and presentation of recommendations improve with AI assistance!

#### <mark style="color:green;">**Quality Enhancement**</mark>

AI assistance enhanced the perceived quality of recommendations across both correct and incorrect answers, with a notable increase in quality scores assigned by human evaluators.

#### <mark style="color:green;">**Differential Effects of AI Treatments**</mark>

The study also observed that the additional overview in the GPT-4 + Overview condition led to better performance compared to the GPT-4 Only group, highlighting the value of guidance in leveraging AI tools effectively.

#### <mark style="color:green;">**Impact on Various Skill Levels**</mark>

The data suggests that AI assistance can enhance the quality of output even when the core recommendation is incorrect, demonstrating a nuanced view of AI's impact on task performance.

In summary, the study illustrates a nuanced landscape of AI's impact on professional work: while AI can significantly enhance productivity and the quality of outputs within its capabilities, its effectiveness diminishes in more complex tasks requiring integrated analytical skills.&#x20;

The findings emphasize the importance of understanding AI's limitations and integrating human oversight, especially in complex decision-making contexts.

### <mark style="color:purple;">Discussion</mark>

#### <mark style="color:green;">**Integration of AI in Knowledge Work**</mark>

* The study highlights AI's dual role: as a productivity and quality enhancer for tasks within its capability frontier and as a potential disruptor for tasks outside this frontier.
* The experiment demonstrated that AI significantly boosts performance and quality within its capabilities, benefiting all workers, especially those with lower initial performance.

#### <mark style="color:green;">**AI as a Booster**</mark>

* AI's assistance led to a notable increase in task speed, quality, and completion rates.
* The study also identified two distinct patterns of AI integration: "Centaurs" who strategically delegate tasks between AI and themselves, and "Cyborgs" who integrate AI deeply into their workflow.

<mark style="color:green;">**AI as a Disruptor**</mark>

* Tasks designed to fall outside AI's frontier showed a decrease in performance when AI was used, highlighting the *<mark style="color:yellow;">**importance of recognising AI's limitations.**</mark>*
* AI's incorrect outputs, when blindly followed, can lead to suboptimal decisions, emphasising the need for critical engagement and validation of AI-generated content.

<mark style="color:green;">**Implications for AI Design and Usage**</mark>

* The findings offer insights for AI tool design, focusing on user navigation and the integration of AI into professional workflows.
* They also prompt discussions on responsible AI usage, especially in high-stakes scenarios, and the need for professional vigilance when working with AI.

<mark style="color:green;">**Organisational and Educational Considerations**</mark>

* The study suggests rethinking how work is organised to better integrate AI, potentially reshaping collaboration, role creation, and adoption strategies within organisations.
* There's a call to maintain a diverse AI ecosystem to avoid idea homogenization and consider the broader competitive landscape where AI's quality enhancements might not always yield distinct advantages.

#### <mark style="color:green;">**Future Directions and Interpretations**</mark>

* The paper concludes with reflections on the transformative potential of AI in high-end knowledge work, comparing its impact on human cognition to the internet's effect on information accessibility.
* It emphasises the ongoing challenge of navigating AI's evolving capabilities and the need for continual adjustment in human-AI collaboration strategies.

Overall, the discussion underscores the complexity of integrating AI into professional settings, highlighting the need for strategic and informed approaches to leverage AI's benefits while mitigating its potential drawbacks.


# Effect of AI on the US labour market

This <mark style="color:blue;">**March 2023**</mark> paper investigates the implications of large language models on the U.S. labour market, focusing on how large language model-powered software enhances capabilities compared to large language models alone.&#x20;

The paper suggests that the impact of LLMs will be substantial, potentially affecting a wide range of occupations and industries.&#x20;

{% embed url="<https://arxiv.org/abs/2303.10130>" %}
An Early Look at the Labour Market Impact Potential of Large Language Models
{% endembed %}

<mark style="color:green;">**Key findings include**</mark>

* About *<mark style="color:yellow;">80% of the U.S. workforce could have at least 10% of their work tasks affected</mark>* by LLMs, with roughly 19% potentially seeing at least 50% of their tasks impacted.
* The effects span all wage levels, suggesting *<mark style="color:yellow;">**higher-income jobs might face greater exposure to LLM capabilities and LLM-powered software**</mark>*.
* The study indicates that LLMs, particularly when enhanced with software and tooling, could *<mark style="color:yellow;">**complete 15% to 56% of all worker tasks more efficiently**</mark>* without sacrificing quality.

### <mark style="color:purple;">Conclusion</mark>

\
This study provides a comprehensive analysis of how Large Language Models (LLMs) might influence various jobs and sectors within the U.S. economy.&#x20;

It introduces a new framework to assess LLM capabilities and their potential impact on employment, revealing that a significant portion of occupations, especially higher-wage ones, are susceptible to LLM integration.&#x20;

About 19% of jobs could see over half of their tasks affected by LLMs, considering both present and forthcoming LLM-enhanced software.

The findings underscore the widespread influence LLMs could have across numerous professions in the U.S., propelled by further advancements in LLM-driven applications.&#x20;

While the technical potential for LLMs to augment or streamline human labour is clear, the actual impact on labour productivity will be shaped by a variety of factors including societal, economic, and regulatory considerations.

As LLM technology progresses, its enduring and expanding effect on the economic landscape will present ongoing challenges for policymakers aiming to foresee and manage its implications.&#x20;

Future research should look into the broader consequences of LLM advancements on the workforce, such as the potential for job augmentation or displacement, effects on job quality, implications for inequality and skill development, and other pertinent issues.&#x20;

Understanding the capabilities and likely impacts of LLMs will enable policymakers and stakeholders to make well-informed choices to steer the complex interplay of AI with the future of work.


# Data Interpreter: An LLM Agent For Data Science

This <mark style="color:blue;">**February 2024**</mark> paper introduces the "<mark style="color:blue;">**Data Interpreter**</mark>," a Large Language Model (LLM)-based agent specifically designed to address the unique and intricate challenges found in data science tasks.&#x20;

The Data Interpreter aims to enhance the problem-solving capabilities of LLMs in scenarios that demand real-time data adjustments, deep optimisation knowledge, and the capacity to identify and correct logical inconsistencies.&#x20;

{% embed url="<https://arxiv.org/abs/2402.18679>" %}
Data Interpreter: An LLM Agent For Data Science
{% endembed %}

### <mark style="color:purple;">Core Objectives of the Data Interpreter</mark>

#### <mark style="color:green;">**Dynamic Planning with Hierarchical Structures**</mark>

The Data Interpreter uses <mark style="color:yellow;">**hierarchical graph structures**</mark> for planning, enabling it to adapt to the dynamic nature of data science tasks.&#x20;

This approach helps the agent understand and navigate the complexities inherent in these tasks, particularly in monitoring data changes and managing dependencies among various variables and processes.

#### <mark style="color:green;">**Tool Integration and Generation**</mark>

The agent enhances its coding proficiency by integrating various human-authored code snippets and creating custom tools for specific tasks.&#x20;

This method goes beyond relying on API calls, allowing the agent to independently build and expand its tool library. This flexibility and self-sufficiency in tool handling enable the Data Interpreter to tailor its approach to each unique problem it encounters.

<mark style="color:green;">**Enhanced Reasoning with Logic Bug Awareness**</mark>

The Data Interpreter is designed to identify logical inconsistencies by using a confidence score derived from execution results and test-driven validations.&#x20;

This feature helps in detecting mismatches between the intended solution and the actual output, allowing for iterative refinement and error reduction in the code it generates.

<figure><img src="/files/wb1Dn09KmRJv0C4uqOag" alt=""><figcaption><p>The overall design of Data Interpreter. This framework consists of three stages: dynamic plan graph and management, wherein a plan is generated for data science problems, and the state of each task is managed during execution; tool utilisation and evolution, involving the selection or creation of suitable tools to solve problems, continually evolving these tools; and automated confidence-based verification, which examines and votes on logically sound solutions.</p></figcaption></figure>

### <mark style="color:purple;">Evaluation and Performance</mark>

The Data Interpreter was evaluated across various data science and real-world tasks, showing notable improvements over existing open-source frameworks.&#x20;

Specifically, it demonstrated a significant increase in performance on machine learning tasks, the MATH dataset, and open-ended tasks. These results underscore the agent's robust problem-solving capabilities and its effectiveness in a wide array of challenges.

### <mark style="color:purple;">Contributions of the Study</mark>

* The paper proposes a novel approach to planning in the context of LLMs, using dynamic and hierarchical structures to enhance adaptability and problem-solving.
* It introduces a method for LLMs to improve their coding abilities through automated tool integration and the generation of custom tools, expanding their capacity to handle diverse and complex tasks.
* It enhances the reasoning capabilities of LLMs by integrating a verification process that improves accuracy and efficiency, addressing one of the critical challenges in deploying LLMs for data science tasks.
* The empirical results provided in the study set new benchmarks for LLM performance in data science, suggesting that the Data Interpreter could serve as a valuable tool for researchers and practitioners in the field.

### <mark style="color:purple;">Related Work</mark>

#### <mark style="color:green;">LLMs as Data Scientist Agents</mark>

The paper highlights how LLMs, trained on a mix of natural and programming languages, have been adapted to handle data science tasks.&#x20;

It mentions several studies where LLMs have been used to decouple complex computations, improve performance on specialised datasets (like the MATH dataset), and enable code-based reasoning in agents.&#x20;

The work also references CodeAct, which dynamically revises code through interactions with a Python interpreter, showcasing an evolving landscape where LLMs are increasingly integrated with code execution to solve data science challenges.

#### <mark style="color:green;">Planning</mark>

In data science, planning involves generating a structured sequence of actions or a roadmap to tackle specific problems.&#x20;

The paper reviews prior works that focus on breaking down complex tasks into smaller, manageable subtasks, and then planning sequentially for these subtasks.&#x20;

It acknowledges the limitations of previous models in handling multi-step problems with strong task dependencies, which are prevalent in data science.&#x20;

To overcome these challenges, the paper introduces a dynamic hierarchical planning approach that allows for more nuanced decomposition of problems into task and action graphs, enhancing the adaptability and efficiency of LLMs in handling complex data science tasks.

#### <mark style="color:green;">Tools</mark>

The section discusses advancements in augmenting LLMs with external tools to enhance their capabilities.&#x20;

Recent studies have focused on not just using tools but also creating and integrating new tools, enabling LLMs to *<mark style="color:yellow;">move from being mere users to creators</mark>*.&#x20;

The paper mentions frameworks that allow LLMs to automatically select and combine tools as needed, which represents a significant step toward more autonomous and versatile AI agents.&#x20;

This shift from static tool assignment to dynamic tool generation and integration reflects a broader trend toward more adaptive and self-sufficient AI systems.

#### <mark style="color:green;">Reasoning</mark>

Reasoning capabilities in LLMs are crucial for processing information and making decisions. The paper reviews works that enhance the reasoning process in LLMs, encouraging them to learn from failures and refine their logic.&#x20;

It discusses pioneering efforts to use code for improving LLMs' accuracy in solving complex mathematical and symbolic reasoning tasks.&#x20;

The paper introduces a novel approach that uses automated confidence-based verification mechanisms to enhance the reasoning capabilities of LLMs, particularly in the context of data science, where advanced logical reasoning is paramount.

### <mark style="color:purple;">Methodology</mark>

In this section, the authors present their methodology for the Data Interpreter. &#x20;

The proposed approach consists of <mark style="color:yellow;">**three main components**</mark>: dynamic planning with a hierarchical structure, tool utilisation and generation, and enhancing reasoning with verification and experience.

<mark style="color:green;">**Dynamic Planning with Hierarchical Structure**</mark>

The authors address the complexity of data science pipelines by organising them using a hierarchical structure.&#x20;

They decompose the problem into manageable tasks and further break down each task into specific actions executed through code. The data science workflows are structured as a hierarchical <mark style="color:blue;">**directed acyclic graph (DAG)**</mark>, representing pipelines at both task and coding levels.

To ensure efficient progress execution and facilitate plan modifications, the Data Interpreter dynamically updates the corresponding code, execution result, and status of each node in the task graph following each execution.&#x20;

The authors introduce two strategies: Self-debugging and Human editing, to enhance autonomous completeness and correctness. If a task fails, Self-debugging utilises LLMs to debug the code based on runtime errors. If the task remains unresolved, Human editing allows for manual modification.

The Data Interpreter regenerates the plan for failed or manually edited tasks based on the current episodic memory and execution context. Throughout execution, the Data Interpreter monitors the dynamic task graph, promptly removing failed tasks, generating refined tasks, and updating the graph.

#### <mark style="color:green;">**Tool Utilisation and Generation**</mark>

To address the intricate nature of tasks that are too complex to be entirely coded from scratch, the authors propose a two-pronged method: tool recommendation and organisation, and continuous tool evolution.

In tool recommendation, the Data Interpreter classifies tools based on task descriptions and types, narrowing down the pool of potential tools. It then identifies the top-k tools that best fit the tasks by evaluating their compatibilities. A tool schema is incorporated to help LLMs understand the functionalities and use cases of these tools.

In tool organisation, LLMs are employed to seamlessly integrate tools into the code, optimally positioning them based on a thorough analysis of the tool functions. The LLM is directed to craft code that invokes the required tool functions and seamlessly integrates these calls with other aspects of the code.

For continuous tool evolution, the Data Interpreter learns from experience during task execution.&#x20;

After each task, it abstracts tools by distilling their core functionalities, creating versatile, generic tool functions that are added to the library for future use. The Data Interpreter automatically ensures the reliability of these tools by conducting rigorous unit tests and leveraging its self-debugging capabilities through LLMs.

#### <mark style="color:green;">**Enhancing Reasoning with Verification and Experience**</mark>

The authors introduce <mark style="color:blue;">**Automated Confidence-based Verification (ACV)**</mark> to evaluate code execution results and determine if the code solution is mathematically rigorous or logically correct.&#x20;

ACV introduces an interpretation layer between the environment and the Data Interpreter.&#x20;

The Data Interpreter generates validation code to ensure that the output result complies with the task requirement. The validation code simulates the logical process according to the task description and verifies the correctness of the result generated by the code.

The Data Interpreter returns a confidence score indicating how likely the output will pass the verification. The confidence score helps the Data Interpreter choose a more accurate result as the final answer by ranking the average confidence scores corresponding to different execution results.

To improve the Data Interpreter's adaptability, the authors integrate an external repository called the 'experience pool' to archive essential elements of each task, including task description, final version code, and final answer.&#x20;

These experiences, including both failed and successful attempts, provide a comprehensive context for a task and can be reused if found to be one of the nearest neighbours of a new task from the vector store.

In summary, the methodology presented in this section combines dynamic planning with a hierarchical structure, tool utilisation and generation, and enhanced reasoning with verification and experience to create an effective LLM-based agent for data science tasks.&#x20;

The approach aims to improve the accuracy, efficiency, and adaptability of the Data Interpreter in handling complex data science problems.

### <mark style="color:purple;">Conclusion</mark>

This paper introduced the Data Interpreter, a solution for data science problem-solving that leverages dynamic planning with hierarchical graphs, tool integration and evolution, and automated confidence-based verification.&#x20;

Through the use of hierarchical graph structures, the Data Interpreter enables efficient decomposition of complex data science problems into manageable tasks and actions. &#x20;

The <mark style="color:blue;">**dynamic planning approach**</mark> ensures real-time adaptability to task variations, allowing for monitoring of data changes and management of intricate variable dependencies. This dynamic nature of the Data Interpreter sets it apart from existing static problem-solving approaches.

The Data Interpreter's <mark style="color:blue;">**tool integration and evolution capabilities**</mark> significantly enhance its coding proficiency and efficiency.  By incorporating human-authored code snippets and creating custom tools tailored to specific tasks, the Data Interpreter continuously expands its toolkit and coding expertise. This adaptive tool utilisation and generation process enables the Data Interpreter to tackle a wide range of data science challenges with improved accuracy and speed.

The <mark style="color:blue;">**automated confidence-based verification mechanism**</mark> introduced in the Data Interpreter further enhances the reliability and reasoning capability of the system.  By evaluating code execution results and determining the logical correctness of code solutions, the Data Interpreter ensures the mathematical rigor and logical soundness of its outputs.  This verification process, coupled with the experience-driven reasoning approach, enables the Data Interpreter to learn from past successes and failures, continuously improving its problem-solving abilities.

Through evaluations on various benchmarks, the Data Interpreter demonstrated superior performance compared to state-of-the-art open-source frameworks.&#x20;

In conclusion, the Data Interpreter represents a significant milestone in the development of LLM-based agents for data science. &#x20;


# The impact of AI on the customer support industry

Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond

This <mark style="color:blue;">**April 2023**</mark> paper explores the impacts of a AI-based conversational assistant on customer support agents' productivity and learning, using data from 5,179 agents.&#x20;

{% embed url="<https://arxiv.org/abs/2304.11771>" %}
Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond
{% endembed %}

### <mark style="color:purple;">Key findings include</mark>

<mark style="color:green;">**Productivity Gains**</mark>

The introduction of the AI tool *<mark style="color:yellow;">**resulted in a 14% increase in productivity across the board**</mark>*.

However, the gains were not uniformly distributed; *<mark style="color:yellow;">**novice and low-skilled workers saw a 34% increase in issues resolved per hour**</mark>*, whereas experienced and high-skilled workers experienced minimal impact.&#x20;

This suggests that the AI tool effectively disseminates best practices and accelerates the learning curve for less experienced workers.

<mark style="color:green;">**Customer Sentiment and Employee Retention**</mark>

Besides boosting productivity, the *<mark style="color:yellow;">**AI assistance improved customer sentiment toward agents and increased employee retention rates**</mark>*, particularly among newer workers.&#x20;

This implies that AI tools can enhance the quality of customer service and the workplace experience for agents.

<mark style="color:green;">**Worker Learning**</mark>

There is suggestive evidence that *<mark style="color:yellow;">**interaction with the AI tool contributes to worker learning**</mark>*.&#x20;

Even during software outages when the AI was unavailable, agents who had more exposure to the AI assistance continued to show improved productivity compared to their pre-AI performance levels.

<mark style="color:green;">**Learning Mechanism**</mark>

The paper posits that *<mark style="color:yellow;">**generative AI systems capture and distribute the tacit knowledge and successful patterns of top-performing agents**</mark>*, making these insights accessible to less experienced or skilled workers.

<mark style="color:green;">**Impact on High-Skill Workers**</mark>

Contrary to the common narrative of skill-biased technical change, where new technologies tend to benefit higher-skilled workers more, this study finds that *<mark style="color:yellow;">**generative AI tools can disproportionately aid less-skilled workers**</mark>*, potentially flattening the productivity distribution within firms.

<mark style="color:green;">**Implications for Future Research**</mark>

The paper calls for further investigation into the broader economic implications of AI adoption, including its effects on wages, labour demand, and skill requirements. It also raises questions about compensating workers for their contributions to the training data that power AI systems.

### <mark style="color:purple;">Economic Impacts</mark>

The economic impacts of generative AI, as discussed in the paper, highlight a significant shift in how computers interact with the workforce.&#x20;

Historically, automation and computerisation have primarily affected routine tasks, leading to a decrease in demand for workers in roles like data entry and an increase in demand for roles requiring complementary skills, such as programming.&#x20;

This shift has contributed to rising wage inequality.

#### <mark style="color:green;">Generative AI tools, however, operate differently</mark>

Generative AI tools do not require explicit instructions to perform tasks but learn from examples, allowing them to undertake activities that involve tacit knowledge, which is knowledge gained through experience and difficult to articulate.&#x20;

This ability *<mark style="color:yellow;">**enables generative AI to perform non-routine tasks that rely on judgment and experience**</mark>*, expanding the scope of tasks that machines can handle and potentially affecting jobs that have been less impacted by previous waves of automation.

However, the deployment of AI in the workplace faces challenges, such as the generation of false or misleading information and the unpredictability of real-world problems compared to controlled laboratory settings.

In addition, the effectiveness of these AI tools in the workplace will likely depend on the interaction with existing organisational structures and may require additional investments or redesigns in business processes.

### <mark style="color:purple;">Customer Support</mark>

The section focuses on the application of Large Language Models (LLMs) for customer support within a Fortune 500 enterprise software company, emphasising the context and potential of generative AI in this domain.

<mark style="color:green;">**Context of Customer Support**</mark>

The customer service industry, marked by high turnover and significant training costs, is increasingly turning to AI tools to enhance efficiency and address workforce challenges. In this setting, customer-agent interactions, critical for maintaining company reputation and customer relationships, vary widely in productivity.

#### <mark style="color:green;">**Generative AI in Customer Support**</mark>

Customer support is an ideal application for generative AI, where conversations can be seen as pattern-matching problems.&#x20;

AI tools can help agents identify and resolve customer issues more effectively by providing real-time suggestions based on past successful and unsuccessful interactions, enhancing the overall customer experience.

#### <mark style="color:green;">**AI Firm and Data Firm Background**</mark>

The study is conducted in collaboration with an AI firm providing AI-based customer support software, focusing on its deployment at a data firm - a Fortune 500 company specializing in business process software.&#x20;

The data firm employs numerous chat-based technical support agents, primarily in the Philippines, to assist U.S.-based small business owners.

Initial observations indicate that agents with AI assistance resolve more chats per hour and exhibit improved performance metrics compared to their pre-treatment and never-treated counterparts. &#x20;

While customer satisfaction remains consistent across groups, post-treatment agents show a notable decrease in average handle time.

The section sets the stage for a detailed examination of how AI deployment affects agent productivity and customer interactions in customer service settings, highlighting the potential for AI to disseminate best practices and improve performance, particularly among newer and less experienced workers.

### <mark style="color:purple;">Conclusion</mark>

The conclusion of the paper emphasises the transformative potential of AI technologies in the workforce, presenting empirical evidence from a real-world setting where a generative AI tool was deployed in customer service. Key findings include:

<mark style="color:green;">**Productivity and Customer Sentiment:**</mark> Access to AI-generated recommendations significantly boosts worker productivity, enhances customer satisfaction, and is linked to lower employee turnover.

<mark style="color:green;">**Dissemination of Best Practices:**</mark> The AI tool appears to capture and distribute the tacit knowledge of high-skill workers, making it accessible to newer and less-skilled employees, thereby democratizing expert knowledge within the organization.

<mark style="color:green;">**Impact on Different Skill Levels**</mark><mark style="color:green;">:</mark> While the AI tool substantially benefits newer and lower-skilled workers by improving their problem resolution and communication styles, it doesn't offer similar advantages to the most skilled or experienced workers.

<mark style="color:green;">**Long-term Implications:**</mark> The study opens up questions about the long-term effects of generative AI on job design, skill requirements, wages, and overall employment in customer service, with potential varying outcomes based on the demand elasticity for customer support.

<mark style="color:green;">**Worker Contribution to AI Training**</mark><mark style="color:green;">:</mark> High-performing workers contribute significantly to the AI's training data but do not see proportional benefits in their productivity, raising questions about appropriate compensation mechanisms for their contributions.

<mark style="color:green;">**Generalisability of Findings**</mark><mark style="color:green;">:</mark> The effects observed in this study might vary in different settings, particularly where product offerings or technical questions are more dynamic, affecting the utility and impact of AI recommendations.

In essence, the study highlights the nuanced and multi-faceted implications of integrating generative AI tools in the workplace, underscoring the need for further research to fully understand their broader economic and organisational impacts.


# Can Large Language Models Reason and Plan?

The answer according to this research is 'no'

The <mark style="color:blue;">**March 2024**</mark> paper examines whether Large Language Models (LLMs) can perform planning and reasoning tasks, traditionally associated with higher cognitive functions.&#x20;

Despite LLMs' impressive linguistic capabilities, the author argues they are essentially sophisticated n-gram models that *<mark style="color:yellow;">**perform approximate retrieval rather than principled reasoning**</mark>*.&#x20;

This paper reinforces our view that generative AI is an augmentation to human work and endeavour, not a replacement.

{% embed url="<https://arxiv.org/abs/2403.04121>" %}
Can Large Language Models Reason and Plan
{% endembed %}

The study tested LLMs like GPT3, GPT3.5, and GPT4 using planning instances from the International Planning Competition and found that while there were improvements in the accuracy of generated plans across versions, the results were <mark style="color:blue;">**not definitive evidence of genuine planning capability**</mark>.

The paper distinguishes between LLMs generating correct responses through memorisation or pattern recognition and performing actual reasoning.  To further test LLMs' planning abilities, the study *<mark style="color:yellow;">**employed obfuscation techniques**</mark>*, which significantly reduced GPT4's performance, suggesting reliance on retrieval rather than planning.

Two methods were explored to potentially enhance LLMs' planning performance: *<mark style="color:yellow;">**fine-tuning with planning data and prompting with hints or external verifiers**</mark>*.  However, fine-tuning didn't show significant improvement, suggesting that it might lead to better approximate retrieval rather than genuine planning.&#x20;

The paper advocates for a framework where *<mark style="color:yellow;">**LLMs' generative capabilities are combined with external verifiers**</mark>* to ensure the correctness and soundness of the planning and reasoning outputs, a setup referred to as <mark style="color:green;">**LLM-Modulo frameworks**</mark><mark style="color:green;">.</mark>

The study concludes that while LLMs exhibit some level of problem-solving capability, their performance in planning and reasoning tasks is largely based on approximate retrieval and memorisation, not genuine reasoning or planning as traditionally understood in AI.

The paper critiques the common practice of iterative prompting by humans in the loop, which may lead to a Clever Hans effect, where the human's input, rather than the LLM's reasoning, guides the outcome. &#x20;

This approach is contrasted with self-improvement methods where LLMs critique and refine their own outputs. However, the author finds such self-verification to potentially worsen performance due to LLMs generating both false positives and negatives.

The author suggests a LLM-Modulo framework *<mark style="color:yellow;">**where LLMs generate potential solutions vetted by external verifiers or expert humans**</mark>*, ensuring a sound outcome. The paper also reflects on the broader implications of LLMs in AI, suggesting they can serve as knowledge sources for domain-specific information, a role previously filled by human knowledge engineers.

In summary, while LLMs demonstrate some level of problem-solving ability, their effectiveness in planning and reasoning is largely attributed to their retrieval capabilities rather than genuine reasoning or planning.&#x20;

The LLM-Modulo framework is proposed as a principled way to leverage LLMs' idea generation for reasoning tasks, supported by external verification to ensure accuracy and soundness.


# KnowAgent: Knowledge-Augmented Planning for LLM-Based Agents

This <mark style="color:blue;">**March 2024**</mark> paper introduces <mark style="color:blue;">**KNOWAGENT**</mark>, a novel framework designed to improve the planning capabilities of language agents, particularly Large Language Models (LLMs).

These agents are integral in AI for tackling complex problem-solving tasks but struggle with sophisticated challenges that require generating executable actions, a limitation attributed to the lack of inherent action knowledge in the models.

{% embed url="<https://arxiv.org/abs/2403.03101>" %}

KNOWAGENT addresses these challenges by *<mark style="color:yellow;">**integrating an external action knowledge base and employing a knowledgeable self-learning strategy.**</mark>*&#x20;

This approach aims to guide the planning process more effectively, enabling the synthesis of more reasonable and coherent action trajectories, thereby enhancing the model's performance in planning tasks.

The framework operates in several steps:

<mark style="color:green;">**Action Knowledge Base Creation**</mark>

An extensive database of action planning knowledge relevant to specific tasks is developed.&#x20;

This knowledge base serves as an external guide for the model's action generation, providing a repository of actions and their corresponding outcomes.

#### <mark style="color:green;">**Knowledge Integration**</mark>

The action knowledge is converted into a text format that the model can understand and use.&#x20;

This integration *<mark style="color:yellow;">allows the model to incorporate external knowledge into its planning process</mark>*, aiding in the generation of more accurate and viable action sequences.

#### <mark style="color:green;">**Knowledgeable Self-Learning**</mark>

The model undergoes a self-improvement phase where it *<mark style="color:yellow;">**refines its understanding and application of the action knowledge through iterative learning**</mark>*. This phase enhances the model's planning accuracy and adaptability.

<figure><img src="/files/84CYG2WXsQ6aLPA9ZyxA" alt=""><figcaption><p>The Path Generation process of KNOWAGENT</p></figcaption></figure>

### <mark style="color:purple;">Background</mark>

The background section elaborates on how language agents model their interaction with the external world, focusing on generating internal thoughts, executable actions, and observing feedback from the environment.

It describes a *<mark style="color:yellow;">**planning trajectory as a series of thoughts, actions, and observations**</mark>*. This sequence helps the agent make decisions and plan its next steps based on previous interactions.

In the KNOWAGENT approach, the paper introduces a sophisticated methodology where the agent uses external action knowledge to enhance its planning capabilities. This method comprises three main steps:

<mark style="color:blue;">**Action Knowledge Definition:**</mark> This part focuses on defining the <mark style="color:yellow;">**action knowledge**</mark> that guides the agent. This knowledge is stored in an action knowledge base, detailing various actions the agent can perform and the associated rules or guidelines for these actions.

<mark style="color:blue;">**Planning Path Generation:**</mark> Using the action knowledge, the agent <mark style="color:yellow;">generates planning paths</mark>. These paths are sequences of actions that the agent could take to achieve its goals. The process involves converting the action knowledge into textual descriptions that the language model can understand and use to formulate plans.

<mark style="color:blue;">**Knowledgeable Self-Learning:**</mark> The agent iteratively refines its planning paths based on the outcomes of its actions. This <mark style="color:yellow;">**self-learning mechanism**</mark> allows the agent to improve its planning capabilities over time, using feedback from its environment and the results of previous plans to make better-informed decisions.

### <mark style="color:purple;">Results</mark>

This section of the paper discusses the experimental setup, results, and analysis of the KNOWAGENT model, which aims to improve the planning capabilities of language agents by integrating explicit action knowledge.&#x20;

<mark style="color:green;">**Main Results**</mark>

* KNOWAGENT consistently outperforms prompt-based methods across different datasets and model sizes.
* The model shows significant improvements over <mark style="color:blue;">**ReAct**</mark>, particularly on the 13b model, highlighting KNOWAGENT's effectiveness in planning path generation.
* Results demonstrate KNOWAGENT's superiority in planning, especially in mitigating planning hallucinations, by adhering to the defined action knowledge.

<details>

<summary><mark style="color:green;">What is the ReAct model?</mark></summary>

The React model, detailed in a paper at ICLR 2023, introduced a novel approach that combined reasoning and acting through language models to enhance task solving in language reasoning and interactive decision-making contexts.&#x20;

The model, referred to as ReAct, interleaves verbal reasoning traces and task-specific actions, enabling dynamic updates to action plans and integration of external information, such as data from APIs like Wikipedia.

</details>

<mark style="color:green;">**Planning Path Generation and Refinement**</mark>

* KNOWAGENT synthesises and refines trajectories using an iterative self-learning process that incorporates action knowledge to filter and merge trajectories, enhancing planning accuracy.
* Ablation studies on action knowledge show that incorporating action knowledge significantly improves model performance and planning quality.

<mark style="color:green;">**Error Analysis**</mark>

* KNOWAGENT shows limitations in handling complex queries and summarising extensive textual data, indicating areas for future improvement in long-text processing and reasoning capabilities.

<mark style="color:green;">**Distilled Knowledge vs. Manually Designed Knowledge**</mark>

* The comparison between manually crafted and distilled action knowledge (from GPT-4) reveals that distilled knowledge is more concise and efficient for simpler tasks.
* For complex tasks requiring longer action sequences, manually designed knowledge outperforms the distilled approach, emphasizing the value of human input in constructing action knowledge.

<mark style="color:green;">**Performance Metrics**</mark>

* The effectiveness of KNOWAGENT is quantified using F1 scores and success rates, with detailed results presented in tables, showing KNOWAGENT's superior performance in planning tasks.

<mark style="color:green;">**Knowledgeable Self-Learning**</mark>

* The iterative fine-tuning process of KNOWAGENT leverages action knowledge to progressively refine the model's planning capabilities, demonstrating the model's ability to learn and improve over iterations.

This detailed analysis highlights KNOWAGENT's innovative approach to enhancing the planning capabilities of language agents by leveraging external action knowledge, demonstrating its effectiveness through comprehensive experiments and analyses.

### <mark style="color:purple;">Conclusion</mark>

KNOWAGENT addresses planning hallucinations by using external action knowledge to inform the generation of synthetic trajectories, enhancing agents' planning proficiency.

The framework employs a self-learning mechanism, translating action knowledge into text for the model's better understanding and utilizing it to guide action generation, demonstrating significant performance improvements over other methods.

Experiments validate KNOWAGENT's efficacy across different models and tasks, establishing its potential in reducing planning errors and enhancing overall agent performance.


# The flaws of 'product-market fit' in an emerging industry

Why is it used as a term and is it relevant in a new industry like artificial intelligence?

Having a strong answer to the question of 'product market fit' is a favourite of the 'start up' scene. &#x20;

Here is one quote:

### ***"Start-ups can't succeed without*** [***achieving product-market fit***](https://posthog.com/blog/product-market-fit-game) ***– it's one of the few things startup gurus agree on."***

We argue this consensus is misplaced in early stage industries - and many start ups are entering these new or quickly evolving markets. &#x20;

Asking an early stage company what their 'product-market' fit is may be apt when the company is entering an existing market with existing competitors with clear strategies and customer segmentation.&#x20;

Investors can work out what the basis of competition in the market is - what the key purchasing factors are within the industry - and make a call on how existing consumers calibrate their purchasing decisions.

They can then identify where they 'fit in' to the existing market.

But we argue <mark style="color:yellow;">p</mark><mark style="color:yellow;">**roduct market fit is not an appropriate metric**</mark> when determining the potential success of a company that is entering <mark style="color:yellow;">an industry going through rapid transformation or evolution.</mark>

### <mark style="color:purple;">How do you define the market?</mark>

Often in early stage markets, the buyers are not aware of the product benefits - there is no identifiable market that can be analysed. &#x20;

So, by definition, in fast-evolving industries, the very act of defining the market itself can be challenging, and the metrics used to gauge product/market fit may not be well-established or may become quickly outdated.&#x20;

This can lead to misconceptions about whether a product truly fits the market or not.

### <mark style="color:purple;">Strategic Flexibility</mark>

Early stage companies in evolving industries must maintain maximum flexibility to ensure they can adapt to market changes.  While this may be unsettling to those that believe in business plans and long term strategic planning, this is the reality.

Often when new industries begin, the rules of the game are unknown, the market is unknown - so why pretend you have  a perfect pre-defined notion of product/market fit?

At best your product/market fit narrative is a hypotheses, that will be tested in the market.  Early stage companies should apply the scientific method - start with a hypothesis, test it, prove it or otherwise, move on or further iterate on the hypothesis.

Product development is an ongoing, iterative process - to suggest you have a well defined 'product market' fit in the early stage of an industry life cycle is a mirage.  &#x20;

You may have a story that could be believable to investors, but forecasting how industries evolve is extremely difficult.

### <mark style="color:purple;">**Evolving Customer Preferences**</mark>

In rapidly changing industries, customer preferences and needs can shift quickly, making it difficult for companies to establish a sustainbale product/market fit.&#x20;

What may seem like a perfect fit today might become obsolete tomorrow as new technologies emerge and consumer behaviour change.

### <mark style="color:purple;">**Competitive Dynamics**</mark>

And even if you achieve 'product-market' fit - what is your sustainable edge?  Why cannot incumbents copy you? &#x20;

Rapidly evolving industries often attract numerous competitors, each trying to establish their own version of product/market fit.&#x20;

This intense competition can lead to rapid changes in the industry landscape, where a company's perceived product/market fit can be disrupted by new entrants or shifts in competitive strategies.

While achieving product/market fit is a significant milestone for any startup, in rapidly evolving industries, it's crucial for companies to remain agile, continuously adapt to changing market conditions, and not rely solely on product/market fit as a measure of success.

The artificial intelligence has many components to it, and each of them are rapidly evolving.  Why pretend that you have a 'product/market' fit, when you do not.


# Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence

The paper "Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence" by Shakked Noy and Whitney Zhang from MIT examines the impact of ChatGPT, a generative AI technology, on the *<mark style="color:yellow;">**productivity of mid-level professionals**</mark>* in performing writing tasks.&#x20;

{% embed url="<https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4375283>" %}
Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence
{% endembed %}

### <mark style="color:purple;">**Abstract and Introduction**</mark>

* The study investigates how ChatGPT affects productivity in professional writing tasks.
* The context is unique due to the generative nature of AI, which differs from traditional automation technologies that focused on routine tasks. This study explores whether generative AI like ChatGPT *<mark style="color:yellow;">**will displace workers or complement them**</mark>*, enhancing productivity.

### <mark style="color:purple;">**Methodology**</mark>

* The experiment involved 444 college-educated professionals who were assigned occupation-specific writing tasks.
* Participants were randomly divided into two groups: one with access to ChatGPT and the other without.
* The tasks were designed to mirror real occupational tasks and included various forms of professional writing.
* Productivity was measured in terms of time taken and the quality of the output, assessed by blinded professionals in the same fields.

### <mark style="color:purple;">**Results**</mark>

* ChatGPT *<mark style="color:yellow;">**significantly increased productivity**</mark>*, reducing the time taken by 0.8 standard deviations and improving output quality by 0.4 standard deviations.
* The *<mark style="color:yellow;">**technology was found to benefit lower-ability workers more**</mark>*, narrowing the productivity gap among workers.
* ChatGPT primarily served as a substitute for worker effort rather than augmenting worker skills, shifting task focus towards idea generation and editing rather than drafting.
* Participants using ChatGPT reported increased job satisfaction and self-efficacy and had mixed feelings of concern and excitement about automation technologies.

### <mark style="color:purple;">**Discussion**</mark>

* The study provides initial insights into how generative AI affects workplace productivity and worker experience.
* The findings suggest that <mark style="color:blue;">**generative AI can enhance productivity and reduce inequalities**</mark> among workers by supporting those with lower abilities.
* The shift in task structure indicates a potential change in the nature of work with the integration of AI technologies.

### <mark style="color:purple;">**Conclusion**</mark>

* The paper concludes that generative AI, specifically ChatGPT, has significant positive effects on productivity and quality in professional writing tasks.
* The results imply that the integration of generative AI in the workplace could lead to a re-evaluation of task allocation and worker roles, particularly in fields involving creative and non-routine tasks.


# The Disruption of the Administrative Class: How Generative AI is Reshaping Organisational Operations

The rise of generative AI is not just a technological advancement; it is a transformative force that threatens to disrupt the foundation of organisational structures.&#x20;

As advanced AI tools offer increasingly sophisticated capabilities in augmenting and, in some cases, replacing traditional administrative functions, the question arises:&#x20;

How will the administrative class adapt to this new reality?

The implications of generative AI extend far beyond the potential replacement of administrators, bureaucrats, and middle management. It strikes at the heart of how organisations operate, *<mark style="color:yellow;">**challenging long-established processes and hierarchies.**</mark>*&#x20;

To fully comprehend the magnitude of this disruption, it is essential to examine the key tasks performed by administrators and assess how artificial intelligence could enhance or replace these functions.

### <mark style="color:purple;">Communication and Coordination</mark>

Effective communication is the lifeblood of any successful organisation.&#x20;

Generative AI models, with their ability to understand and process natural language, are poised to revolutionise this domain.  Acting as advanced communication coordinators, these models can draft emails, schedule meetings, and facilitate seamless inter-departmental coordination.&#x20;

Their capacity to handle complex queries, redirect communications to appropriate departments, and provide swift responses to stakeholders promises to streamline organisational communication like never before.

### <mark style="color:purple;">Resource Management</mark>

In the resource management process, generative AI brings efficiency and accuracy.&#x20;

By analysing financial data and forecasting expenditures, these models can assist in budgeting processes, ensuring optimal allocation of resources.

In procurement, AI can automate vendor communications, manage purchase orders, and maintain transaction records, streamlining the entire procurement lifecycle.  The potential for cost savings and improved decision-making in resource management is immense.

### <mark style="color:purple;">Planning and Scheduling</mark>

Planning and scheduling is where fine-tuned AI models truly shine.

With their ability to manage calendars, set reminders, and organise events, these models can take on the bulk of administrative planning tasks. Their integration with various scheduling tools ensures optimal allocation of time and resources, aligning with organisational priorities.&#x20;

The days of manual scheduling and coordination may soon be a memory.

### <mark style="color:purple;">Record Keeping and Documentation</mark>

Generative AI is set to redefine record-keeping and documentation.&#x20;

With their ability to transcribe meetings, generate reports, and maintain records with accuracy and consistency, these models can streamline administrative processes.&#x20;

The volume of data that AI can process and organise ensures that all records are updated, categorised, and easily retrievable, reducing the burden on human administrators.

### <mark style="color:purple;">Policy and Procedure Compliance</mark>

Navigating the complex world of laws and regulations is a critical responsibility of administrators.&#x20;

Generative AI models, programmed to understand and stay current with relevant legal standards, can be a game-changer in compliance management.&#x20;

By reviewing organisational policies, suggesting necessary updates, and ensuring that all practices adhere to the latest legal requirements, AI can help organisations stay compliant and mitigate risks.

### <mark style="color:purple;">Problem-Solving and Decision Making</mark>

While generative AI models cannot fully replace human judgment, their ability to analyse historical data and current trends can provide invaluable insights and recommendations.&#x20;

By supporting administrators in making informed decisions, AI can enhance problem-solving capabilities and drive more effective decision-making processes.

### <mark style="color:purple;">Supporting Executive Functions</mark>

Administrators often serve as the backbone of executive support, managing correspondence, organising travel arrangements, and preparing materials for meetings and presentations.&#x20;

Generative AI models can take on many of these tasks, allowing executives to focus on more strategic aspects of their roles.  The potential for increased efficiency and productivity in executive support is immense.

As the integration of generative AI models in administrative functions becomes increasingly prevalent, it heralds a new era of efficiency and effectiveness in organisational management.&#x20;

By automating routine tasks, providing analytical insights, and enhancing communication, these models empower administrators to focus on strategic decision-making and innovation.

However, the reality is that, over time, *<mark style="color:yellow;">**the administrative class may find itself increasingly replaced by generative AI.**</mark>*&#x20;

The disruptive potential of these technologies cannot be ignored, and organisations must be prepared for a future where many traditional administrative roles are automated.

### <mark style="color:purple;">What to do?</mark>

The path forward for the administrative class lies in adaptation and upskilling.&#x20;

By embracing AI as a tool for augmentation rather than replacement, administrators can position themselves as strategic partners in the AI-driven workforce.&#x20;

Those who can leverage AI to enhance their decision-making capabilities, provide valuable insights, and drive innovation will find themselves at the forefront of this transformative era.

The disruption brought about by generative AI is not a distant future; it is a present reality.&#x20;

Organisations that fail to recognise and adapt to this shift risk being left behind in an increasingly competitive landscape.&#x20;

The administrative class must rise to the challenge, embracing the opportunities presented by AI while navigating the complexities of this new paradigm.


# How Knowledge Workers Think Generative AI Will (Not) Transform Their Industries

{% embed url="<https://arxiv.org/abs/2310.06778>" %}


# Embracing AI: A Strategic Imperative for Modern Leadership

The adoption of AI is not just a technological upgrade but a <mark style="color:yellow;">required implementation</mark> for leadership in various sectors.&#x20;

Successful AI implementation demands leaders who understand both the technical and business aspects of AI.   Leaders must balance technical feasibility with business impact to ensure the viability and success of AI projects.

The complexity of AI projects often leads to a high failure rate, financial losses, and scepticism about AI capabilities.&#x20;

#### <mark style="color:green;">Role of Chief AI Officer</mark>

Appointing a Chief AI Officer bridges the technical and business aspects, highlighting AI's strategic importance. This role is crucial for guiding AI initiatives and ensuring alignment with business objectives.

Understanding these challenges is crucial for effective implementation and management.

### <mark style="color:purple;">Early Adoption and Potential Growth</mark>

Despite AI becoming increasingly familiar, many large companies are still in the nascent stages of adoption. This presents a unique opportunity for forward-thinking leaders to explore and harness its potential to gain a competitive edge.

### <mark style="color:purple;">Excitement and Experimentation</mark>

There's excitement among enterprises about using AI to enhance products and transform business operations. From automating customer service to content creation, and from code writing to troubleshooting, the applications of AI in customer support, sales, marketing, and engineering are vast and varied.

### <mark style="color:purple;">Diverse Application Scenarios</mark>

Enterprises are identifying numerous use cases for AI across different departments and functions. The challenge lies in determining which of these use cases are immediately feasible and valuable.

### <mark style="color:purple;">Main Use Case Categories</mark>

The use cases for AI in enterprises broadly fall into information discovery and synthesis, hierarchical summarisation, and support/chatbots.  These categories encompass the majority of current AI applications in the enterprise sector.

### <mark style="color:purple;">The Road Ahead for Leadership - General Notes</mark>

<mark style="color:green;">**Enterprise Knowledge Integration:**</mark> Leaders should focus on creating AI models that integrate and make internal knowledge accessible while maintaining governance and security.

<mark style="color:green;">**Adoption by Non-Tech Enterprises:**</mark> There is significant potential for enterprises using older technologies to leapfrog to the forefront of AI adoption.

<mark style="color:green;">**CIO Interest:**</mark> CIOs are key drivers in adopting AI. Developing solutions that appeal to the CIO’s broad perspective can be a strategic move.

<mark style="color:green;">**Compliance and Security:**</mark> Addressing data governance, security, and compliance is crucial, especially for enterprises hesitant to share data with external parties.

<mark style="color:green;">**Addressing AI Challenges:**</mark> Tackle issues like hallucination in models and improve attribution to ensure reliability and transparency.

<mark style="color:green;">**Data Privacy and Cost-Quality Trade-off:**</mark> Address data privacy concerns effectively and offer models that balance cost and quality.

<mark style="color:green;">**Building Trust Through Reference Customers:**</mark> Acquire early adopters and showcase their success stories to build trust among potential customers.

<mark style="color:green;">**Rethinking Language Model Architecture:**</mark> Innovate in how AI models are structured, focusing on reducing hallucination and improving attribution.

<mark style="color:green;">**Compliance and Legal Concerns:**</mark> Be proactive in addressing potential legal issues and ensure models comply with legal standards.

<mark style="color:green;">**Focus on Use Cases Over Technology:**</mark> Centre discussions with clients around specific use cases and problems they are trying to solve.

<mark style="color:green;">**Enhanced Access to Knowledge:**</mark> Generative AI can solve the challenge of knowledge sharing in consulting, providing comprehensive insights.

<mark style="color:green;">**AI Governance and Use Case Exploration:**</mark> Implement a structured approach to exploring AI's potential, as seen in PwC’s establishment of a governance group and an AI factory.

<mark style="color:green;">**Improving Customer Experience with AI:**</mark> Leverage AI to rethink customer engagement and personalise interactions in a non-intrusive manner.

<mark style="color:green;">**Efficiency and Quality Gains:**</mark> Generative AI is expected to significantly improve efficiency across various business processes.

<mark style="color:green;">**Resource Reallocation and Cost Reduction:**</mark> AI offers the potential to reduce operating costs, allowing for resource reallocation towards innovation.

<mark style="color:green;">**Marketing Challenges:**</mark> There is a risk of overemphasising AI's strategic value without addressing practical implementation aspects, leading to a disparity between expectations and reality.  A balanced approach to marketing AI's potential and challenges is essential.

<mark style="color:green;">**Infrastructure and Cultural Shifts:**</mark> AI requires not just specialised compute infrastructure but also a cultural shift within organisations.  Embracing rapid innovation and disruption is key to successful AI adoption.

<mark style="color:green;">**Adoption Challenges:**</mark> Building trust between humans and AI is critical. The uncertainty inherent in AI adds complexity, making trust a crucial factor for successful adoption. Overcoming this challenge involves transparent communication and showcasing AI's reliability and benefits.

<mark style="color:green;">**Transformative Potential of AI:**</mark> Despite the high failure rates, the successful 20% of AI projects can revolutionise business operations. AI's ability to remove human bottlenecks and enable scalability is evident from an array of implementations.

<mark style="color:green;">**Quantifying AI Value:**</mark>  Many companies struggle to quantify the value added by AI, often overstating its effectiveness. Establishing concrete metrics and benchmarks is crucial for accurately assessing AI's impact.

<mark style="color:green;">**Need for New Infrastructure and Culture:**</mark> Significant AI value leaps require not only new technological infrastructure but also a fundamental cultural shift. Organizations must embrace rapid innovation and disruptive changes to fully leverage AI's potential.

<mark style="color:green;">**Cultural and Process Transformation**</mark><mark style="color:green;">:</mark> AI integration involves more than just technology; it requires a shift towards data-driven decision-making and redefining business processes.

Traditional methods become obsolete, and models like the 'AI factory' centralise data and streamline AI development.

<mark style="color:green;">**AI Development Process:**</mark> AI development processes often resemble pre-industrial workflows, marked by inefficiency and error-proneness. A structured and efficient process is essential for successful AI deployment.

<mark style="color:green;">**Experimentation and Adaptation:**</mark> Staying competitive with AI technologies requires experimentation and leadership engagement. Understanding and investing in AI is critical for realising its full potential.

<mark style="color:green;">**Regulatory and Ethical Considerations:**</mark> The fast-paced growth of AI, particularly generative AI, creates regulatory uncertainties and ethical concerns. Building trust in AI requires transparency and ethical algorithm development. Additionally, AI audits for fairness and bias are becoming increasingly important.

### <mark style="color:purple;">Conclusion</mark>

Integrating AI into business operations demands a balanced approach encompassing technical expertise, strategic vision, ethical considerations, and a deep understanding of AI technologies.

C-suite executives must navigate these complexities thoughtfully, leveraging a phased approach and a framework for trust to harness AI's potential effectively. The future of business lies in embracing AI's transformative power while responsibly managing its challenges.

***


# Artificial Intelligence and Management: The Automation-Augmentation Paradox

The complex interplay between automation and augmentation in the use of artificial intelligence (AI) in management

This <mark style="color:blue;">**2023**</mark> paper explores the dual roles of artificial intelligence (AI) in the area of management—automation and augmentation.

It examines the prevailing narrative encouraged by recent business literature, which advises organisations to prioritise augmentation, where humans collaborate with AI, over automation, where AI replaces human tasks.

{% embed url="<https://archive-ouverte.unige.ch/unige:169300>" %}
RAISCH, Sebastian, KRAKOWSKI, Sebastian. Artificial Intelligence and Management: The Automation Augmentation Paradox
{% endembed %}

### <mark style="color:purple;">Key points from the paper include</mark>

<mark style="color:green;">**Automation vs. Augmentation**</mark>

Automation refers to AI systems taking over tasks previously performed by humans, while augmentation involves AI assisting humans in tasks, enhancing their capabilities.

<mark style="color:green;">**Interdependence of Automation and Augmentation**</mark>

The authors challenge the clear-cut separation between automation and augmentation presented in business literature.  They argue that these two roles of AI are interconnected and create a paradoxical tension within organisational and management contexts.

<mark style="color:green;">**Paradoxical Tension**</mark>

This tension arises because focusing too much on either automation or augmentation can lead to negative consequences for both organisations and society.

Overemphasis on automation may lead to job losses and societal issues, whereas overemphasis on augmentation may ignore the efficiencies and benefits automation can offer.

<mark style="color:green;">**AI's Impact on Society and Business**</mark>

The paper underscores the broader implications of AI's dual roles, highlighting the need for thoughtful consideration of how AI strategies impact not just organisational performance but also societal outcomes.

### <mark style="color:purple;">Some Insights</mark>

The analysis expands on the persistence of the tension between automation and augmentation in AI's application within management, emphasising that this tension is paradoxical because it's enduring and involves interdependent yet contradictory elements.&#x20;

<mark style="color:green;">**Persistence of Tension:**</mark> The tension between automation and augmentation in management is enduring due to the inherent limitations and distinctive roles of machines and humans.&#x20;

Despite advancements in AI, *<mark style="color:yellow;">**machines cannot fully replace human intelligence**</mark>*, particularly in complex managerial tasks, underscoring the ongoing necessity for human involvement.

<mark style="color:green;">**Machine Limitations**</mark>

Several intrinsic limitations of machines highlight the need for ongoing human interaction:

* Machines lack self-awareness and purpose, requiring humans to define objectives and take responsibility.
* In complex tasks, machines provide options, but human intuition is necessary for final decision-making.
* AI systems are trained for specific tasks and can't generalize their learning to unrelated domains.
* Machines lack human sensory perceptions, emotions, and social skills, essential in nuanced managerial tasks.

<mark style="color:green;">**Cyclical Relationship**</mark>

Automation and augmentation have a cyclical relationship where initial automation might lead to further augmentation in adjacent tasks, and vice versa.&#x20;

Over time, tasks that were augmented may become automated as understanding and technologies evolve, but changing conditions may necessitate a return to augmentation.

<mark style="color:green;">**Management Strategies**</mark>

Organisations may fall into vicious cycles if they focus narrowly on either automation or augmentation, neglecting the interplay between the two. &#x20;

This can lead to detrimental effects, such as deskillment, complacency, and lock-in to automated processes, or continual failure and escalating commitment in augmentation efforts.

<mark style="color:green;">**Virtuous Cycles**</mark>

A constructive approach involves recognising the paradoxical nature of the tension and adopting strategies that combine differentiation (leveraging the distinct benefits of automation and augmentation separately) and integration (iterating between automation and augmentation to exploit their respective strengths).

This balanced approach can foster innovation, adaptability, and comprehensive engagement with AI in management.

<mark style="color:green;">**Organisational Outcomes**</mark>

* Augmentation can significantly boost productivity, enhance service quality, and spur innovation by merging human intuition with machine efficiency.
* Differentiation between automation and augmentation allows organisations to reap unique benefits from each: automation brings cost efficiency and consistency, while augmentation fosters creativity and adaptability.
* Integrating automation and augmentation can lead to synergies, like freeing up resources through automation for more complex, augmented tasks, potentially enabling innovative business models like personalised medicine.

<mark style="color:green;">**Societal Outcomes**</mark>

* The paradox has broad implications beyond individual organisations, affecting labour markets and social equality.
* Focusing solely on automation might lead to job losses and increased unemployment, exacerbating social inequality.  Conversely, an exclusive focus on augmentation might deepen the digital divide, creating disparities between those who can and cannot engage in augmented tasks.
* Balancing automation and augmentation could foster a cycle of deskilling in areas where machines excel and upskilling in areas where human skills are paramount, potentially enhancing job satisfaction by shifting focus to more creative and fulfilling tasks.
* The use of AI in management could impact social equality and fairness. Automation might reduce human biases in decisions, promoting fairness, while augmentation could help mitigate machine biases through human oversight.

In conclusion, navigating the automation-augmentation paradox effectively requires a nuanced approach that acknowledges the interdependence of these two aspects of AI in management.&#x20;

Organisations that successfully balance and integrate automation and augmentation can not only enhance their performance and innovation but also contribute positively to broader societal challenges like employment and social justice.


# Network effects in AI models

The Role of Artificial Intelligence and Data Network Effects for Creating User Value

The highly cited paper "The Role of Artificial Intelligence and Data Network Effects for Creating User Value" by Robert Wayne Gregory et al. investigates the impact of artificial intelligence (AI) and data network effects on the perceived value of platforms for users.&#x20;

The authors introduce the concept of data network effects, *<mark style="color:yellow;">**where a platform becomes more valuable as it learns more from user data**</mark>*, enhancing the user experience through AI-driven personalisation and improvement.

{% embed url="<https://www.semanticscholar.org/paper/The-Role-of-Artificial-Intelligence-and-Data-for-Gregory-Henfridsson/c6918eafcca310643b1efe7eddcb779477f6d87f>" %}
The Role of Artificial Intelligence and Data Network Effects for Creating User Value" by Robert Wayne Gregory et al
{% endembed %}

### <mark style="color:purple;">**Introduction and Background**</mark>

* The paper discusses how network effects have traditionally been a significant factor in the value users perceive in platforms.   Network effects occur when a product or service becomes more valuable as more people use it.
* Traditionally, research has focused on direct network effects (value from user interactions) and indirect network effects (value from complementary products or services).
* The authors propose a new category called <mark style="color:green;">**data network effects**</mark>, where the value increases as the platform learns from user data, enhancing personalisation and service quality.

### <mark style="color:purple;">**Artificial Intelligence**</mark>

* AI is explained through three principles: <mark style="color:green;">**combination**</mark> (integrating various technologies), <mark style="color:green;">**recursiveness**</mark> (interdependent modular architecture), and <mark style="color:green;">**phenomena**</mark> (focus on data-driven learning).
* The evolution of AI is attributed to advancements in machine learning, improved computing power, and the ability to process and manage vast amounts of data.
* *<mark style="color:yellow;">**AI's value lies in its ability to continuously learn and improve from data**</mark>*, which in turn enhances user experience and value perception.

### <mark style="color:purple;">**Data Network Effects**</mark>

* The authors theorise that the AI capability of a platform contributes to data network effects by enabling the platform to learn from user data and improve its services or products for each user.
* This learning leads to improvements in product functionality, platform quality, and user experience, thereby increasing the platform's value.
* *<mark style="color:yellow;">**Data network effects are presented as a new form of network externality**</mark>* where a user's utility from a platform is influenced by the platform's AI-driven data learning and improvements.

### <mark style="color:purple;">**Model Development**</mark>

* The paper develops a model to explain how AI capabilities and data network effects contribute to creating user value, extending existing network effects theory.
* The model suggests that AI's role in platforms can lead to significant enhancements in user value through continuous learning and service improvement based on user data.

### <mark style="color:purple;">**Discussion**</mark>

* The paper discusses the implications of data network effects for platform strategies and user engagement.
* It highlights the importance of AI in enabling platforms to leverage data network effects for competitive advantage and enhanced user value.

### <mark style="color:purple;">**Conclusion**</mark>

* By introducing the concept of data network effects, the paper expands the understanding of how network effects contribute to user value.
* It emphasises the pivotal role of AI in harnessing the power of user data to enhance platform value, offering a new perspective on the interplay between AI, data, and user value in digital platforms.

### <mark style="color:purple;">Data network effects</mark>&#x20;

Traditionally, network effects are understood through direct and indirect interactions among users, where the utility a user derives from a platform grows with the network's size.&#x20;

However, the integration of AI introduces a novel dimension to this concept: data network effects.

#### <mark style="color:green;">**Data Network Effects**</mark>

AI transforms platforms by using data network effects, where the value to users increases as the platform leverages data through AI for learning and improvement.  This marks a shift from the traditional focus solely on network size to also include the quality of interactions and enhancements driven by data analysis.

#### <mark style="color:green;">**AI's Role**</mark>

AI facilitates scaling of data-driven learning from users' digital interactions, influencing the platform's value by providing personalised experiences or improved functionalities.&#x20;

This learning leads to new platform externalities, impacting user value beyond just the network size.

#### <mark style="color:green;">**Transaction Feasibility and Engagement**</mark>

The paper highlights how network structure and user engagement affect transaction feasibility and user value.&#x20;

Platforms like Uber use AI to optimise matchings between supply and demand, engaging users more effectively and enhancing the platform's value.

#### <mark style="color:green;">**Network Conduct and User Behaviour**</mark>

Network conduct, influenced by AI, affects user value by managing and moderating user interactions to prevent opportunistic behaviour and promote trustworthiness. Platforms use AI to maintain a reliable environment, encouraging positive user experiences and interactions.

#### <mark style="color:green;">**Trust and Reputation**</mark>

The perception of trust and reputation within the platform is crucial for user engagement and value creation. AI algorithms help build and maintain trust, facilitating smoother transactions and interactions among users.

Overall, the integration of AI into platforms introduces data network effects as a new paradigm in understanding value creation, emphasizing the role of data-driven learning and user engagement in enhancing user value.

### <mark style="color:purple;">Conceptual Framework</mark>

The framework section of the paper outlines a comprehensive approach to understanding how artificial intelligence (AI) and data network effects enhance user value on platforms.&#x20;

#### <mark style="color:green;">**Platform AI Capability**</mark>

AI enables platforms to learn from data, enhancing their ability to predict and improve services for users.&#x20;

This capability is crucial for creating value, as it allows platforms to adapt and offer personalised experiences based on user data.  The paper argues that the *<mark style="color:yellow;">**speed and accuracy of AI-driven predictions significantly impact perceived user value**</mark>*, making platforms more responsive and tailored to user needs.

#### <mark style="color:green;">**Data Stewardship**</mark>

The quantity and quality of data play a vital role in training machine learning algorithms. &#x20;

Higher data quantity allows for more robust and diverse training, enhancing the model's predictive power.  Similarly, *<mark style="color:yellow;">**high-quality data ensure that predictions are reliable and relevant, boosting user value**</mark>*.  Effective data stewardship ensures that the AI has the right data to learn from, enhancing its ability to create value.

#### <mark style="color:green;">**User-Centric Design**</mark>

The design of platform services and products influences how users interact with and perceive AI capabilities.&#x20;

User-centric design focuses on understanding and meeting user needs, which encourages engagement and enables users to experience the benefits of AI directly. Performance and effort expectancy are highlighted as crucial factors; *<mark style="color:yellow;">**platforms must be designed to meet user expectations and be easy to use to foster engagement and enhance value creation**</mark>*.

#### <mark style="color:green;">**Platform Legitimation**</mark>

The paper emphasises the importance of platform legitimation, which involves ensuring that the *<mark style="color:yellow;">**platform's use of data and AI is perceived as ethical and appropriate by stakeholders**</mark>*.&#x20;

This includes addressing concerns about data privacy, security, and the explainability of AI decisions. By aligning with stakeholder expectations and norms, platforms can maintain access to essential resources and support, further enhancing their ability to create user value.

Through these components, the framework illustrates a holistic view of how AI and data network effects interact with various platform aspects to enhance user value, emphasising the interconnectedness of technology, data, design, and ethical considerations in the digital platform ecosystem.

### <mark style="color:purple;">Discussion</mark>

The discussion highlights the nuanced interdependencies between various types of network effects and AI capabilities in shaping user value, suggesting a multifaceted approach to understanding platform dynamics in the AI era.&#x20;

It also outlines potential avenues for future research, including exploring the interplay of artificial and collective intelligence and examining the implications of data network effects on traditional competitive dynamics.


# AI impact on the publishing industry

Will AI take over from human creativity? &#x20;

The prospect of AI generated content replacing human creativity is abhorrent to most - regardless, we are likely to see an evolving dynamic between humans and AI in creative endeavours.&#x20;

We have already seen the AI creativity controversy through one case study in the book:

### &#x20;<mark style="color:green;">**'Death of an Author' by Stephen Marche**</mark>

This novel highlights the potential of AI-assisted writing - the author admitting that 95% of the content was generated by AI.

The book is a typical detective story, and while it was not critically acclaimed - it has been described by one critic as 'not awful'!

This book has caused some consternation.  We are starting to witness a plethora of AI generated book content on major platforms like Amazon.  And there has been a backlash.  Self-publishers must now declare if content sold on Amazon’s site is AI-generated

The debate will continue about the interaction of human creativity and AI.   But the publishing industry should be looking to apply the technology in a positive fashion - seeking to complement and enhance the creative process - as well as market lesser known authors and Indie books.

### <mark style="color:purple;">AI Driven Search Engines</mark>

The current process of searching for a book is likely to change for the better. &#x20;

The current process of filtering by fiction or non-fiction, genres such as romance, horror or crime, or by author is restrictive.  Furthermore, there is well known circular feedback loop - where books that rate well on customer reviews, continue to rate well just because they have the most purchases. &#x20;

#### <mark style="color:green;">Does this stifle up and coming authors?</mark>  &#x20;

With AI increasingly driving search this process will likely change.  AI models will better understand user preferences and align book recommendations to them.  &#x20;

Publishing houses need to think more about how they adapt.  A generative AI model can be fine tuned to optimise book metadata and descriptions for AI-enhanced search engines.  This will ensure better visibility in a crowded digital landscape.

### <mark style="color:purple;">Personalised Stories</mark>

The publishing industry is experiencing an increasing number of submitted manuscripts, but support staff to help manage these submissions has been reduced, increasing the burden on editors.

AI could be used here, to assist determine what could be considered a story that could sell. &#x20;

Customised neural language models can assist publishers in the manuscript selection process by quickly analysing submissions and identifying those with the most potential, reducing the workload of editorial teams.

Generative AI can be fine tuned to develop personalised reading experiences, such as <mark style="color:yellow;">custom short stories or tailored book suggestions based on the reader’s preferences and reading history</mark>.

With the evolving nature of social media platforms, there's an opportunity to create models that help authors adapt their content and marketing strategies to these emerging platforms, maximising their reach and engagement with newer audiences.

### <mark style="color:purple;">Content Creation</mark>

Generative AI integrated with knowledge bases can assist author researching specific topics - in thematic research, character development, and plot construction. This tool could generate ideas, suggest reading materials, or even help in drafting story elements based on the author's input.

This assistant could also provide suggestions for improving writing style, grammar, and coherence.

### <mark style="color:purple;">Marketing</mark>

Decades of corporate consolidation in the publishing industry have led to fewer opportunities for advancement and a more corporate, less creative work environment. This consolidation also impacts the variety of books being acquired, with a focus on more commercially viable titles.

Generative AI models can be customised and tailored for authors and publishers to generate marketing copy, ad creatives, and promotional materials.     These tools could help in creating compelling blurbs, social media posts, and ad banners, saving time and enhancing marketing efforts.

### <mark style="color:purple;">Indie Authors</mark>

An indie book is one that is independently published, usually by the author, without the involvement of a traditional publishing house. &#x20;

Generative AI can support indie work by providing market analysis, sales trend predictions, and personalised business advice based on the author's genre and publishing history.

These tools could also assist in planning and executing effective book launches - including timeline planning, task management, promotional material generation, and performance analytics.

### <mark style="color:purple;">AI for Reader Community Engagement</mark>

An LLM application could be *<mark style="color:yellow;">**developed to facilitate deeper interaction with readers**</mark>*, perhaps through AI-driven Q\&A sessions, personalised responses to reader queries, or interactive book discussions.

### <mark style="color:purple;">Conclusion</mark>

The publishing industry is on the cusp of a significant transformation, driven by the rapid advancements in generative AI.&#x20;

While the prospect of AI-generated content replacing human creativity may seem daunting, the reality is that AI will more likely complement and enhance the creative process.&#x20;

By embracing AI technologies, publishers can streamline their operations, discover new talent, and create more engaging and personalised experiences for readers.&#x20;

Authors, too, can benefit from AI-assisted tools that aid in research, writing, and marketing. As the industry navigates this new landscape, it is crucial to find a balance between leveraging AI's capabilities and preserving the unique human touch that makes literature so compelling.

### <mark style="color:purple;">AI Application: "BookMate" - Your AI-Powered Publishing Assistant</mark>

BookMate is an innovative AI-powered platform designed to revolutionise the publishing industry by assisting authors, publishers, and readers throughout the book creation and consumption process.&#x20;

<mark style="color:green;">**This comprehensive tool combines several key features**</mark>

<mark style="color:blue;">**Manuscript Evaluation:**</mark> BookMate uses customized neural language models to analyze submitted manuscripts quickly, identifying those with the highest potential and reducing the workload of editorial teams.

<mark style="color:blue;">**Author Assistance:**</mark> The platform provides a generative AI-powered writing assistant that helps authors with thematic research, character development, plot construction, and writing style improvements.

<mark style="color:blue;">**Personalised Reader Experiences:**</mark> BookMate employs generative AI to create tailored book suggestions and even custom short stories based on readers' preferences and reading history.

<mark style="color:blue;">**Marketing Support:**</mark> The platform offers a suite of AI-driven marketing tools that generate compelling blurbs, social media posts, and ad creatives, helping authors and publishers promote their books more effectively.

<mark style="color:blue;">**Indie Author Support:**</mark> BookMate provides market analysis, sales trend predictions, and personalised business advice to support indie authors in their publishing journey.

<mark style="color:blue;">**Reader Engagement:**</mark> The platform facilitates deeper interaction between authors and readers through AI-driven Q\&A sessions, personalized responses to reader queries, and interactive book discussions.

By integrating these features into a single, user-friendly platform, BookMate aims to empower authors, publishers, and readers alike, fostering creativity, discoverability, and engagement in the ever-evolving world of publishing.


# Power asymmetry

The concept of power asymmetry pervades various aspects of society, manifesting in relationships between employers and employees, consumers and corporations, students and educational systems, and across the international stage among states.&#x20;

These imbalances often lead to disparities in negotiations, rights, opportunities, and outcomes for the less powerful parties involved.&#x20;

However, the advent of generative AI holds the potential to democratise access to information, resources, and platforms, *<mark style="color:yellow;">**potentially reducing these power asymmetries**</mark>*.&#x20;

By providing all individuals and entities with advanced tools for knowledge generation, decision-making support, and enhanced communication capabilities, generative AI should be able to <mark style="color:green;">**level the playing field**</mark>**.**&#x20;

### <mark style="color:purple;">Academic Work on Power Asymmetry</mark>

**"Power and Interdependence" by Robert O. Keohane and Joseph S. Nye**

This seminal work in international relations theory discusses how states and actors manoeuvre in a world where power is distributed unevenly. It introduces concepts like soft power and complex interdependence.

**"Negotiating Power: Agenda Setting and the Asymmetry of Influence" by Deborah M. Kolb and Judith Williams**

This study in the context of negotiations explores how power asymmetries affect the setting of agendas and outcomes in negotiation processes.

**"Power in Organizational Societies" by Michael Mann**

Mann's work delves into the role of power in social and organisational structures, discussing its sources and effects on societal dynamics.

**"The Power Paradox: How We Gain and Lose Influence" by Dacher Keltner**

This book explores how power is acquired, how it can corrupt, and how it influences social behavior. It provides insights into the psychological aspects of power dynamics.

**"The Bases of Social Power" by John R.P. French and Bertram Raven**

This foundational study in social psychology categorizes different bases of power, such as coercive, reward, legitimate, referent, and expert power. It's widely referenced in understanding power dynamics in various contexts.


# Information Asymmetry

Applied AI can reduce this major economic inefficiency

Information asymmetry in decision-making occurs when there is an imbalance in the amount of information available to the parties involved.&#x20;

This imbalance can lead to market inefficiencies, as the party with more information might manipulate or distort decisions to their advantage.&#x20;

### <mark style="color:purple;">What are the ramifications of information asymmetry?</mark>

The primary issue is a misallocation of resources and and an increase in costs. &#x20;

<mark style="color:green;">**Increased Transaction Costs**</mark><mark style="color:green;">:</mark> To counteract information asymmetry, entities may incur additional costs in the form of due diligence, audits, or insurance, which can elevate the overall cost of doing business.

<mark style="color:green;">**Market Failures**</mark><mark style="color:green;">:</mark> In extreme cases, information asymmetry can lead to market failures, where the market does not operate efficiently or effectively, potentially causing significant economic disruptions.

<mark style="color:green;">**Loss of Trust**</mark><mark style="color:green;">:</mark> When parties cannot trust the information they receive, it erodes confidence in markets and institutions.

<mark style="color:green;">**Inequality**</mark><mark style="color:green;">:</mark> Information asymmetry can exacerbate inequality. Those with better access to information can exploit their advantage, often at the expense of those who are less informed, leading to a concentration of wealth.

<mark style="color:green;">**Moral Hazard and Adverse Selection**</mark><mark style="color:green;">:</mark> These phenomena can lead to risky behaviors that would not occur if information were symmetric, potentially leading to significant societal costs, such as financial crises or environmental degradation.

<mark style="color:green;">**Exploitative Practices and Fraud**</mark><mark style="color:green;">:</mark> Information asymmetry can enable unethical practices and fraud, leading to individual and collective harm, eroding trust in systems and institutions.

<mark style="color:green;">**Regulatory Costs**</mark><mark style="color:green;">:</mark> Governments may need to intervene to correct information asymmetries, leading to regulatory costs and the potential for regulatory capture or inefficiencies.

### <mark style="color:purple;">Generative AI to the rescue?</mark>

However, the advent of generative AI presents a potential solution to this issue.

By analysing extensive datasets and uncovering hidden insights through big data analytics and deep learning, AI can make decisions with fewer biases and limitations than humans.&#x20;

Consequently, the deployment of AI is anticipated to reduce information asymmetry in various markets.&#x20;

Enhanced AI models can provide more balanced information across different sectors, *<mark style="color:yellow;">**fostering more efficient and equitable decision-making processe**</mark>*&#x73;. This technological advancement is expected to significantly mitigate the adverse effects of information asymmetry and the resulting power imbalances in markets.

### <mark style="color:purple;">Here are examples of where generative AI models could reduce asymmetry</mark>

<mark style="color:green;">**Healthcare Services**</mark>

Information asymmetry in healthcare is prevalent, with patients often lacking the knowledge or resources to understand their health conditions fully or the implications of their treatment options.&#x20;

Domain focused generative AI models could assist by digesting vast amounts of medical literature and patient data to provide personalised, understandable, and accurate health information to patients, enabling them to make more informed decisions about their care.

<mark style="color:green;">**Real Estate**</mark>

The real estate market often involves significant information asymmetry between buyers, sellers, and intermediaries.&#x20;

Domain focused generative AI models could help reduce this by analysing market trends, property histories, neighbourhood data, and regulatory information to provide both buyers and sellers with comprehensive insights into property values, investment potential, and legal considerations, thus leveling the playing field.

<mark style="color:green;">**Insurance**</mark>

The insurance industry is characterised by asymmetric information between insurers and policyholders, particularly in terms of risk assessment and policy terms.&#x20;

Generative AI models could play a role in demystifying policy documents and claims processes for consumers while helping insurers more accurately assess risk by analysing vast datasets on claims history, environmental factors, and individual behavior patterns.

<mark style="color:green;">**Automotive Sales**</mark>

Information asymmetry in the automotive market can lead to issues like lemon markets, where sellers have more information about the vehicle's condition than buyers.&#x20;

Generative AI could be used to aggregate and analyse data from vehicle history reports, maintenance records, and user reviews to provide potential buyers with a comprehensive understanding of a vehicle's condition and history.

<mark style="color:green;">**Car Maintenance and Repairs**</mark>

Often, vehicle owners lack the technical knowledge to understand the specifics of their car's maintenance needs or the extent of repairs required when issues arise.&#x20;

This gap can lead to distrust or overcharging by service providers.&#x20;

An generative AI model could help bridge this gap by offering vehicle owners personalised maintenance advice, explanations of common car issues in layman's terms, and fair price estimates for repair works based on a vast database of car models, local service rates, and user reviews.

<mark style="color:green;">**Plumbing Services**</mark>

Homeowners typically have limited knowledge about plumbing, making it challenging to diagnose issues or estimate the cost of repairs accurately. An AI assistant could provide homeowners with preliminary diagnostics based on symptoms described, suggest possible solutions, and offer cost estimates to empower them during negotiations with service providers.

<mark style="color:green;">**Mental Health Services**</mark>

In the field of psychology and psychiatry, patients often struggle to find the right specialist for their specific needs due to a lack of understanding of the different specialties within mental health services.&#x20;

An AI platform could help reduce this asymmetry by matching patients with therapists or psychiatrists based on detailed analysis of their conditions, treatment history, and preferences, while also educating them about different therapeutic approaches.

<mark style="color:green;">**Legal Services**</mark>

The complexity of legal language and concepts can create information asymmetry between legal professionals and their clients.&#x20;

An AI tool could translate legal jargon into plain language, offer basic legal advice based on case law and statutes, and help clients understand their rights and options. This would not replace legal professionals but could empower individuals to make more informed decisions about pursuing legal action.

<mark style="color:green;">**Home Renovation and Construction**</mark>

For many homeowners, undertaking renovation or construction projects involves navigating a maze of building codes, material options, and design choices, often without the expertise to make informed decisions.&#x20;

An generative AI model could analyse project goals, budget constraints, and local regulations to offer tailored advice on materials, design choices, and cost-saving opportunities, reducing dependence on contractors for information.

<mark style="color:green;">**Educational Services and Tutoring**</mark>

Students and parents often struggle to identify the most suitable educational paths or tutoring services due to a lack of detailed understanding of a student’s strengths, weaknesses, and learning styles.&#x20;

A generative AI model could analyse a student's performance data, learning habits, and preferences to recommend personalised educational programs, courses, and tutoring services.&#x20;

This would ensure that students receive guidance that is closely aligned with their needs, thereby enhancing learning outcomes.

<mark style="color:green;">**Financial Planning and Investment Advice**</mark>

Many individuals find financial planning and investment daunting due to the complex nature of financial products and the uncertainty of market conditions.&#x20;

AI could demystify this by providing personalised financial advice based on an individual's financial goals, risk tolerance, and market trends.&#x20;

<mark style="color:green;">**Travel Planning and Booking**</mark>

Planning a trip involves sifting through vast amounts of information about destinations, accommodations, activities, and logistics, often leading to decision fatigue and suboptimal choices.&#x20;

An generative AI travel assistant could alleviate this by offering personalised travel recommendations based on the traveler’s interests, budget, travel history, and reviews from similar travelers.

It could also optimise itineraries by analysing travel times, costs, and local attractions, ensuring a tailored and enjoyable travel experience with minimal effort from the traveler.

### <mark style="color:purple;">Conclusion</mark>

AI has the potential to transform markets by reducing information asymmetry, leading to a decrease in trade volume but an increase in market efficiency.

It advocates for a future where markets are increasingly dominated by AI agents, hypothesizing a shift towards more rational and efficient economic environments.


# Continuum

Applied Artificial Intelligence

## <mark style="color:purple;">Overview</mark>

Continuum's foundation is the application of technology.  &#x20;

We translate the latest academic research and technology into practical application.

This knowledge repository is an overview of the research and technology of each of the primary components of a neural language model application and the generative AI industry.

We decompose the <mark style="color:blue;">**generative AI application value chain**</mark> into the following sub components:

1. Datasets - Creation, Curation and Structuring
2. Models - Foundation Models, Instruction Tuned Models
3. Fine Tuning - Tokenization,  Embedding, Parameter Efficient Fine Tuning, Training Processes
4. Inference - Optimisation of model inference
5. Knowledge - Vector databases
6. Retrieval Augmented Generation
7. Embedding and Recommendation Engines
8. AI agents and AI autonomy
9. Regulation and Ethics
10. AI Infrastructure

### <mark style="color:purple;">Generative AI Disruption</mark>

We also use this platform to discuss the potential disruption of generative AI on various domains such as data architecture and pipelines, search and recommendation systems.

Continuum Labs is an organisation that seeks to participate in the disruption that will be caused by artificial intelligence.

We offer a suite of tools and a platform designed for the development of generative AI applications - seeking to empower organisations of all sizes to harness the disruptive potential of AI - and concurrently avoid being disrupted by AI.

Our mission is to not only provide the technical foundation for AI applications, but to generate creative ideas that can be implemented to secure competitive edge for businesses.

Our approach goes beyond offering tools; we aim to provide a perspective on how AI can create organisational value.&#x20;

We see generative AI disrupting major industries - content creation, advertising, knowledge sectors, consulting, creatives and most important of all - search. &#x20;

With this will come the total reorganisation of data architecture.  Generative AI will become embedded in data pipelines, redirecting data flows, transforming it, and translating it into knowledge and insights to allow enhanced and more rapid decision making.

## <mark style="color:purple;">Quick links</mark>

{% content-ref url="/pages/lkGiHdGvoygkVk231tq3" %}
[Datasets](/data/datasets)
{% endcontent-ref %}

{% content-ref url="/pages/BIkKGigNzyPAQp6beObX" %}
[MODELS](/models/foundation-models)
{% endcontent-ref %}

{% content-ref url="/pages/J7U0QPLW4AwLJxwCyf12" %}
[Training](/training/the-fine-tuning-process)
{% endcontent-ref %}

{% content-ref url="/pages/Z4X7N8yXkfkRUrT6OgRa" %}
[INFERENCE](/inference/why-is-inference-important)
{% endcontent-ref %}

{% content-ref url="/pages/Tkl8QsbL8ufiPOAN4xoq" %}
[Vector Databases](/knowledge/vector-databases)
{% endcontent-ref %}

{% content-ref url="/pages/MFNPgly3lfrotMKSoJjn" %}
[Retrieval Augmented Generation](/knowledge/retrieval-augmented-generation)
{% endcontent-ref %}

{% content-ref url="/pages/veoOcxTDKaNoy3uxZZi0" %}
[AGENTS](/agents/what-is-agency)
{% endcontent-ref %}

{% content-ref url="/pages/35b0RU2bpapniMWWMVM4" %}
[Regulation and Ethics](/regulation-and-ethics/regulation-and-ethics)
{% endcontent-ref %}

{% content-ref url="/pages/iAfrDPmZYxgkeySlXxHn" %}
[DISRUPTION](/disruption/data-architecture)
{% endcontent-ref %}

{% content-ref url="/pages/P6J6Cd1yOazv1oF4oPKk" %}
[Search](/disruption/search)
{% endcontent-ref %}

{% content-ref url="/pages/GUE8KF62OXkjyjwowuVB" %}
[Infrastructure](/infrastructure/the-modern-data-centre)
{% endcontent-ref %}


# Datasets

The data used to train neural language models has always been important, but as the industry has evolved it has learnt that <mark style="color:green;">more data is not always better</mark>.

### <mark style="color:purple;">Data Quantity and Scaling Laws</mark>

The relationship between model size, training dataset size, and data repetition has historically been considered crucial.  Research illustrates that model performance can be systematically improved by scaling up model size and training data size concurrently.&#x20;

However, recent research has show that quality of data, and how it is structured and ingested into the model training process is just as important. &#x20;

### <mark style="color:purple;">Ensuring Data Quality</mark>

We have found quality assurance techniques are indispensable during pretraining and fine tuning phases.&#x20;

Practices like deduplication, quality filtering, and toxicity filtering serve multiple purposes: they enhance training efficiency, minimise privacy risks, and reduce model memorisation.&#x20;

The focus on deduplication is particularly noteworthy, as it prevents train-test overlap, improving model perplexity and reliability.

### <mark style="color:purple;">Domain Composition and Data Diversity</mark>

The diversity and domain composition of datasets are paramount.

Creating heterogeneous dataset compositions ensures models are equipped with a broad range of abilities - reflecting the expanse of knowledge. Techniques for domain re-weighting and composition highlight the necessity for datasets to be balanced and varied, underscoring the importance of inclusivity in training data.

But this is just the case for creating 'general models' - creating models for highly specific use cases does not necessarily require this level of diversity.

This section on data explores the history of datasets used for trainiing foundation models, as well as those use for fine tuning pre-trained models.


# Pre Training Data

Training Foundation Models

Foundation models are the foundation of the generative AI industry.   Post the release of a paper that led to the development of the Transformer architecture, foundation model development has continued to grow.

The first public release of a pre-trained large language model based on the Transformer was the 117 million parameter GPT model by OpenAI.  Following this, various models of increasing size were developed by many different companies including OpenAI, Google, Meta, Microsoft and Nvidia.

Foundation models are trained on extensive corpora, often encompassing billions of documents.  Most of this text data has been derived from public sources. but there is an increasing demand for proprietary or private datasets.

The capacity of these models to accurately predict subsequent elements in a sequence is based on their understanding of language patterns and context learnt from the datasets on which they are trained.

These models have been the bedrock of the generative AI revolution, offering both proprietary and increasingly open-source options for a range of applications.&#x20;

Their development over recent years has underscored the potential of large-scale models to transform various sectors by providing advanced capabilities in language understanding and generation.

### <mark style="color:purple;">Sources of Training Data</mark>

#### <mark style="color:green;">The Internet</mark>

The Internet has been a rich resource for pre-training LLMs, offering a breadth of linguistic knowledge due to the wide range of content available online.&#x20;

However, the quality of such data varies greatly, with high-quality sources like Wikipedia and lower-quality sources such as spam emails. &#x20;

This source of data will continue to remain important, but it is critical this data is cleaned to improve its quality.

#### <mark style="color:green;">**Conversational Text**</mark>

Conversational text from sources like Reddit or social media platforms is used to improve model ability to engage in dialogues and perform on question-answering tasks.&#x20;

The primary issue with this source of data is the invasion of privacy.  Conversational text, while on public forums - was never conceived to be used as training data for artificial intelligence.

#### <mark style="color:green;">**Books**</mark>

Book data provides formal and coherent long texts, contributing to a model's ability to understand complex linguistic structures and dependencies.

Open-source datasets like Books3 and Bookcorpus2, found in the Pile dataset, are common sources for this type of data, which aids LLMs in generating narrative texts and understanding formal language.

### <mark style="color:purple;">The Pile</mark>

The most famous source of training data for foundation models is known as the "The Pile".

The Pile is a 825 gigabyte diverse, open source language modelling data set that consists of 22 smaller, high-quality datasets combined together.  The Pile&#x20;

{% embed url="<https://pile.eleuther.ai/>" %}
The Pile
{% endembed %}

### <mark style="color:purple;">The primary constituents of "The Pile"</mark>

<table><thead><tr><th width="202">Source</th><th width="145" align="center">% of Dataset</th><th>Details</th></tr></thead><tbody><tr><td>Pile-CC</td><td align="center">18.11%</td><td>Common Crawl: Collection of website crawls, including web pages, metadata, and text extractions.</td></tr><tr><td>PubMedCentral</td><td align="center">14.40%</td><td>PubMed Central: Subset of PubMed repository for biomedical articles, with open full-text access.</td></tr><tr><td>Books3†</td><td align="center">12.07%</td><td>Dataset of books derived from Bibliotik private tracker, mix of fiction and non-fiction.</td></tr><tr><td>OpenWebText2</td><td align="center">10.01%</td><td>OpenWebText2: Web scraped dataset inspired by WebText and OpenWebTextCorpus, with content from Reddit.</td></tr><tr><td>ArXiv</td><td align="center">8.96%</td><td>Preprint server for research papers predominantly in Math, Computer Science, and Physics.</td></tr><tr><td>Github</td><td align="center">7.59%</td><td> Large corpus of open-source code repositories, enabling code-related task improvements.</td></tr><tr><td>FreeLaw</td><td align="center">6.12%</td><td>Access to and analytical tools for academic studies in the legal realm from the FreeLaw Project.</td></tr><tr><td>StackExchange</td><td align="center">5.13%</td><td>User-contributed content on a network of question-answer websites covering various subjects.</td></tr><tr><td>USPTOBackgrounds</td><td align="center">3.65%</td><td>Background sections from patents granted by the US Patent and Trademark Office.</td></tr><tr><td>PubMedAbstracts</td><td align="center">3.07%</td><td>Abstracts from publications in PubMed, covering a wide range of biomedical topics.</td></tr><tr><td>Gutenberg (PG-19)†</td><td align="center">2.17%</td><td>Classic Western literature from Project Gutenberg books pre-1919.</td></tr><tr><td>OpenSubtitles</td><td align="center">1.55%</td><td>English dataset of subtitles from movies and TV shows, providing natural dialog</td></tr></tbody></table>

{% embed url="<https://arxiv.org/abs/2101.00027>" %}
The paper with details of "The Pile"
{% endembed %}


# Types of Fine Tuning

Some definitions

### <mark style="color:purple;">Supervised Fine-Tuning</mark>

Supervised is when <mark style="color:yellow;">you tell the model what the answer is</mark> - the most basic type

* **Definition**: Supervised fine-tuning involves training a model on a <mark style="color:yellow;">labeled dataset,</mark> where each piece of data (like a sentence or image) is paired with a correct answer or label.
* **Process**: The model learns by comparing its predictions to these correct answers and adjusting itself to improve accuracy.
* **Example**: Training a language model to translate sentences by using a dataset where each English sentence is paired with its French translation.

### <mark style="color:purple;">Unsupervised Fine-Tuning</mark>

* **Definition**: Unsupervised fine-tuning does not use labeled data. Instead, the <mark style="color:yellow;">model learns patterns or features from the data itself</mark> without explicit guidance on what is correct.
* **Process**: The model tries to find structure in the data, like grouping similar things together or predicting parts of the data based on other parts.
* **Example**: Training a model to group news articles into topics without knowing in advance what the topics are.

### <mark style="color:purple;">Self-Supervised Fine-Tuning</mark>

The answer comes naturally from the task

* **Definition**: Self-supervised learning is a <mark style="color:yellow;">hybrid approach that falls between supervised and unsupervised learning</mark>. It involves creating a <mark style="color:yellow;">pseudo-labeled dataset from the unlabeled data.</mark>
* **Process**: The model itself generates labels from the data (hence "self-supervised") and then trains on these labels. Often, this involves masking parts of the data and training the model to predict them.
* **Example**: A language model is given sentences with some words missing and learns to predict the missing words.

### <mark style="color:purple;">Summary</mark>

* **Supervised**: Needs labeled data and learns to predict correct answers.
* **Unsupervised**: No labels are involved, and the model learns patterns or structures from the data itself.
* **Self-Supervised**: Generates its own labels from the data and learns like in supervised learning, but without needing external labels.

Each of these methods has its own strengths and is suitable for different types of problems and datasets.&#x20;

Supervised learning is very direct and effective when labeled data is available, unsupervised learning is useful for uncovering hidden structures in data, and self-supervised learning offers a balance by exploiting unlabeled data in a structured way.


# Self Instruct Paper

The most highly cited paper on fine tuning methods

Language models that are fine-tuned to follow human-written instructions have shown remarkable abilities in understanding and generating text.&#x20;

However, they face limitations due to their dependence on a limited amount of human-written instruction data, which lacks diversity and creativity.  These constraints hinder the model's ability to generalise across a wider range of tasks.

To address these limitations, this important <mark style="color:blue;">**May 2023**</mark> paper introduced the <mark style="color:blue;">SELF-INSTRUCT</mark> framework.

This framework uses a bootstrapping approach, where the language model generates its own instruction, input, and output samples.&#x20;

These generated samples are then refined and used to fine-tune the original model. This approach creates an almost annotation-free method for aligning pre-trained language models with instructions, overcoming the constraints posed by limited human-written instruction data.

{% embed url="<https://arxiv.org/abs/2212.10560>" %}
Self-Instruct Paper
{% endembed %}

### <mark style="color:purple;">The Limitation of Current Instruction-Tuned Models</mark>

At the core of traditional instruction-tuned models lies their dependency on human-written instructions.  This dependency creates a bottleneck, limiting the quantity, diversity, and creativity of instruction data available for model training.&#x20;

As a result, the models' ability to generalise and perform across a broad spectrum of tasks is constrained.&#x20;

### <mark style="color:purple;">Introducing SELF-INSTRUCT: A Paradigm Shift</mark>

The SELF-INSTRUCT framework emerged as a solution to overcome the limitations of traditional instruction-tuned models.&#x20;

At its heart, SELF-INSTRUCT employs a bootstrapping method that enables the language model to generate its own instruction, input, and output samples.&#x20;

This  approach not only minimises the need for human-annotated data but also introduces a higher level of diversity and creativity in the instruction data generated.&#x20;

The generated samples are then pruned and used to fine-tune the original model, aligning it more closely with human-written instructions while significantly reducing the dependency on human-generated content.

### <mark style="color:purple;">Best Practices for Creating Self-Instruct Datasets</mark>

Creating effective self-instruct datasets involves a combination of strategic planning, iterative development, and diverse inputs. Here are some best practices to consider:

#### <mark style="color:blue;">**Diverse and Representative Seed Instructions**</mark>

* **Goal**: Ensure the initial seed instructions cover a broad spectrum of tasks across different domains to promote a wide-ranging dataset.
* **Example**: Starting with seeds that include instructions for culinary recipes, technical troubleshooting, academic essay writing, and fitness exercise guides.

#### <mark style="color:blue;">**Iterative Refinement**</mark>

* **Goal**: Continuously improve the quality of the dataset by generating instructions, assessing their utility and clarity, and refining them based on feedback.
* **Example**: Using an initial dataset to train a model, then using the model’s output to identify gaps and errors which guide the creation of revised instructions that better meet the needs of the tasks.

#### <mark style="color:blue;">**Quality Control Mechanisms**</mark>

* **Goal**: Implement processes to remove or revise instructions that are unclear, biased, or overly simplistic.
* **Example**: Employing automated tools to detect and eliminate biased language or using manual reviews to ensure instructions meet a set of quality standards.

#### <mark style="color:blue;">**Balancing Novelty and Coherence**</mark>

* **Goal**: Create instructions that are both innovative and logical, ensuring they lead to executable tasks.
* **Example**: Developing instructions for a new game or app functionality that are understandable and actionable by a broad user base.

### <mark style="color:purple;">Examples of Self-Instruct Datasets</mark>

To illustrate these best practices, here are examples of datasets created following the above guidelines:

#### <mark style="color:blue;">Dataset for Common Daily Tasks</mark>

* **Instruction**: Plan a daily schedule.
* **Input**: List of tasks: Doctor's appointment at 9 AM, Team meeting at 11 AM, Gym at 3 PM, Grocery shopping at 5 PM.
* **Output**: "Your daily schedule is: 9 AM - Doctor's appointment, 11 AM - Team meeting, 1 PM - Lunch break, 3 PM - Gym, 5 PM - Grocery shopping."

#### <mark style="color:blue;">Dataset for Text Processing Tasks</mark>

* **Instruction**: Summarise the following article.
* **Input**: Text: "Climate change is increasingly seen as an urgent global issue, with rising temperatures affecting diverse ecosystems."
* **Output**: "The article highlights the urgent need to address global warming and its impacts on ecosystems."

#### <mark style="color:blue;">Dataset for Educational Content Creation</mark>

* **Instruction**: Create a quiz question based on the given information.
* **Input**: Topic: "Photosynthesis", Details: "Photosynthesis is the process by which plants convert sunlight into energy, using water and carbon dioxide, releasing oxygen as a byproduct."
* **Output**: "What are the essential elements involved in photosynthesis and what byproduct does it release?"

#### <mark style="color:blue;">Dataset for Code Generation Tasks</mark>

* **Instruction**: Write a Python function to calculate the Body Mass Index (BMI).
* **Input**: Null
* **Output**:

```python
def calculate_bmi(weight, height):
    return weight / (height ** 2)
```

### <mark style="color:purple;">Empirical Evidence of Performance Gains</mark>

The efficacy of SELF-INSTRUCT is not just theoretical.  When applied to models like GPT-3, the framework has demonstrated substantial improvements.

Specifically, it achieved a 33% performance boost on the SUPERNATURALINSTRUCTIONS dataset, nearly matching the performance of InstructGPT, which benefits from private user data and human annotations.&#x20;

Human evaluators have also confirmed that models fine-tuned with SELF-INSTRUCT surpass those tuned with existing public instruction datasets, marking a significant leap forward in model performance.

### <mark style="color:purple;">Beyond Performance: Democratising Instruction-Based Fine-Tuning</mark>

SELF-INSTRUCT's impact extends beyond performance metrics.&#x20;

By minimising the reliance on human annotations, the framework democratises the process of instruction-based fine-tuning.

This  is particularly important given the previously noted challenges with the scalability and generalisability of instruction-following models due to the reliance on human-annotated data.&#x20;

SELF-INSTRUCT's approach also opens the door to exploring its application in commercial settings, particularly in automating or semi-automating the fine-tuning process for bespoke applications.

### <mark style="color:purple;">A New Horizon for Research</mark>

The introduction of SELF-INSTRUCT represented a  shift in how we approach the fine-tuning of language models.&#x20;

By automating the generation of diverse and creative instruction data, the framework addresses the critical bottlenecks of human annotation and the limitations of existing public datasets.&#x20;

Furthermore, the SELF-INSTRUCT framework has potential applications in multi-modal learning, indicating its versatility and the broad implications of its use.


# Self-Alignment with Instruction Backtranslation

Xian Li, Ping Yu, Chunting Zhou, Timo Schick, Omer Levy, Luke Zettlemoyer, Jason Weston, Mike Lewis

In this <mark style="color:blue;">March 2024</mark> paper, the authors introduce a novel method called "instruction backtranslation" to create high-quality instruction-following language models without relying on large amounts of human-annotated data. The approach leverages a small amount of seed data and a large web corpus to automatically generate and curate training examples.

{% embed url="<https://arxiv.org/abs/2308.06259>" %}
Self-Alignment with Instruction Backtranslation
{% endembed %}

The key steps of the instruction backtranslation method are as follows:

<mark style="color:green;">Self-augmentation:</mark> The seed model generates instruction prompts for web documents, creating potential training examples.

<mark style="color:green;">Self-curation:</mark> The seed model selects high-quality examples from the generated candidates.

<mark style="color:green;">Fine-tuning:</mark> The selected high-quality examples are used to fine-tune a stronger model.

<mark style="color:green;">Iteration:</mark> The process is repeated, using the improved model to better curate the instruction data and re-train the model.

The authors highlight the importance of data quality in aligning large language models (LLMs) for instruction following.&#x20;

While human-annotated datasets are valuable, they are difficult to scale. The instruction backtranslation method addresses this challenge by leveraging the model itself to augment and curate training examples, enabling self-alignment.

The approach draws inspiration from the <mark style="color:blue;">backtranslation method in machine translation</mark>, where target sentences are automatically annotated with model-generated source sentences in another language. In this case, the model generates instruction prompts for web documents and selects high-quality (instruction, output) pairs for training.

The authors demonstrate the effectiveness of their approach by fine-tuning LLaMa on two iterations of instruction backtranslation.  The resulting model, named Humpback, outperforms all other non-distilled models on the Alpaca leaderboard, showcasing the power of self-alignment through iterative self-augmentation and self-curation.

### <mark style="color:purple;">Method</mark>

The instruction backtranslation method consists of two main steps: self-augmentation and self-curation. The process is iterative, allowing the model to improve its ability to select high-quality examples for fine-tuning. Let's break down each step in detail:

<mark style="color:green;">Initialization</mark>

* Start with a base language model (e.g., LLaMa), a small seed dataset of human-annotated (instruction, output) pairs, and a large unlabeled web corpus.
* Preprocess the web corpus by extracting self-contained segments, deduplicating, filtering by length, and removing low-quality segments.

<mark style="color:green;">Self-Augmentation</mark>

* Fine-tune the base language model on (output, instruction) pairs from the seed data to create a backward model Myx, which predicts instructions given outputs.
* For each unlabeled example yi in the web corpus, use the backward model to generate a candidate instruction ˆxi.
* Create candidate augmented paired data A := {(ˆxi, yi)} by combining the generated instructions with their corresponding outputs.

<mark style="color:green;">Self-Curation</mark>

* Start with a seed instruction model M0 fine-tuned on (instruction, output) pairs from the seed data.
* Use M0 to score each augmented example (ˆxi, yi) in A and derive a quality score ai using prompting (e.g., instructing the model to rate the quality on a 5-point scale).
* Select a subset of the augmented examples with scores ai ≥ k to form a curated set A(1)k.

<mark style="color:green;">Iterative Self-Curation</mark>

* Use the curated augmentation data A(t-1)k from the previous iteration, along with the seed data, to fine-tune an improved model Mt.
* Use Mt to rescore the augmented examples for quality, resulting in a new augmentation set A(t)k.
* Perform multiple iterations of data selection and fine-tuning to obtain the final model (e.g., M2 after two iterations).
* When combining seed data and augmented data for fine-tuning, use tagging to distinguish the data sources (e.g., append "Answer in the style of an AI Assistant." for seed data and "Answer with knowledge from web search." for augmented data).

<mark style="color:green;">Example of emulating the process</mark>

1. Start with a base model like GPT-3 and a small seed dataset of human-annotated (instruction, output) pairs, along with a large web corpus like Common Crawl.
2. Fine-tune GPT-3 on (output, instruction) pairs from the seed data to create a backward model that predicts instructions given outputs.
3. For each document in the web corpus, extract self-contained segments and use the backward model to generate candidate instructions for each segment.
4. Create candidate augmented paired data by combining the generated instructions with their corresponding segments.
5. Fine-tune a seed instruction model (e.g., GPT-3) on the (instruction, output) pairs from the seed data.
6. Use the seed instruction model to score each augmented example using prompting (e.g., "On a scale of 1 to 5, how well does the output answer the given instruction?").
7. Select a subset of the augmented examples with scores above a certain threshold (e.g., 4 or 5) to form a curated set.
8. Fine-tune the seed instruction model on the curated set, along with the seed data, to create an improved model.
9. Use the improved model to rescore the augmented examples and create a new curated set.
10. Repeat steps 8 and 9 for multiple iterations to obtain the final instruction-following model.

By emulating this process, you can leverage large amounts of unlabeled data to create high-quality instruction-following models without relying heavily on human annotation.

### <mark style="color:purple;">Experiments</mark>

The experiments in this paper aim to evaluate the effectiveness of the proposed instruction backtranslation method for training instruction-following language models. The authors conducted several experiments to analyze the impact of data quality, data quantity, and various ablations. Let's break down the experiments in detail:

#### <mark style="color:green;">Experimental Setup</mark>

* Seed data: 3,200 high-quality (instruction, output) pairs from the first turn of the Open Assistant dataset.
* Base model: LLaMA with 7B, 33B, and 65B parameters, fine-tuned using the same hyperparameters as existing supervised fine-tuning methods.
* Unlabeled data: 502k segments from the English portion of the Clueweb corpus.
* Baselines: text-davinci-003, LIMA, and Guanaco.
* Evaluation: 1,130 unique prompts from various sources, with a dev set of 256 prompts. Automatic evaluation using AlpacaEval and human preference evaluation.

#### <mark style="color:green;">Seed and Augmentation Data Statistics</mark>

* Analysis of instruction and output lengths for seed data, self-augmented data, and self-curated data.
* Task diversity analysis using the verb-noun structure of instructions.

#### <mark style="color:green;">Scaling Analysis</mark>

* Data quality vs. data quantity: Fine-tuning on augmented data of different quality (without curation, A(2)4, and A(2)5) to understand the importance of data quality.
* Data scaling efficiency: Comparing the performance of various instruction-following models as the amount of fine-tuning data changes. Estimating the scaling coefficient α for different instruction datasets.

#### <mark style="color:green;">Model Quality</mark>

* AlpacaEval: Evaluating the generation quality using GPT-4 as the judge, comparing Humpback to non-distilled, distilled, and proprietary models.
* Human Evaluation: Pairwise comparison of Humpback with open-source and proprietary models using human preference judgments.
* Commonsense Reasoning and MMLU: Zero-shot accuracy on five commonsense reasoning benchmarks and the Massive Multitask Language Understanding (MMLU) benchmark.

#### <mark style="color:green;">Ablations</mark>

* Training on self-augmented data only: Comparing the performance of models trained on self-augmented data with and without self-curation, and jointly fine-tuning with seed data.
* System prompts: Analyzing the effects of using system prompts to distinguish augmented data from seed data during fine-tuning and inference.

The experiments demonstrate that the proposed instruction backtranslation method, using self-augmentation and self-curation, can effectively leverage large amounts of unlabeled data to create high-quality instruction-following models. The results show that Humpback outperforms other non-distilled models and achieves competitive performance compared to distilled and proprietary models. The ablation studies further confirm the importance of self-curation and the complementary nature of seed data and augmented data.

### <mark style="color:purple;">Related Work</mark>

The related work section discusses various approaches to instruction tuning for large language models (LLMs) and the challenges in gathering high-quality demonstration examples for fine-tuning.&#x20;

Early work on instruction tuning focused on NLP tasks, showing that fine-tuning with instruction-output pairs improves cross-task generalization.  Recent work has extended instruction tuning to a broader range of general tasks, incorporating instructions from LLM users.

Existing high-quality instruction-following LLMs rely on human annotations, which are expensive and time-consuming to collect. Some works have explored using LLMs to generate instructions, such as Unnatural Instructions, Self-Instruct, and the concurrent work by Köksal et al. (2023). However, these approaches either use model-generated responses for training data or rely on distillation from a more powerful model.

The authors also discuss self-alignment, where the model is utilized to improve itself and align its responses with desired behaviors. Many of these works construct training data in an unsupervised way or use the model to generate additional context to condition on at inference time.

The importance of data quality is highlighted, with approaches like PALMS and LIMA showing that curating high-quality human-written data results in strong performance. The concurrent work by Chen et al. (2023) provides an algorithmic approach to select high-quality data.

Finally, the authors discuss distillation, where most fine-tuned LLaMA models are based on knowledge distillation from ChatGPT or GPT-4. These approaches require an already strong model and do not provide a recipe for building a strong model from scratch.

### <mark style="color:purple;">Conclusion</mark>

In conclusion, the proposed instruction backtranslation method offers a scalable approach to fine-tuning large language models for instruction following. By leveraging large amounts of unlabeled data and using the model itself to augment and curate high-quality training examples, this iterative self-training algorithm enables the creation of strong instruction-following models without relying heavily on human annotations or distillation from more powerful models.

The experiments demonstrate that the Humpback models, fine-tuned using instruction backtranslation, outperform all other non-distilled instruction-following models on the Alpaca leaderboard while using fewer human-annotated examples. This showcases the effectiveness of the self-augmentation and self-curation steps in improving the model's performance.

The analysis suggests that scaling this method further by considering larger unlabeled corpora could yield even greater gains. As the field of instruction tuning for LLMs continues to evolve, the instruction backtranslation approach presents a promising direction for creating high-quality, general-purpose instruction-following models in a more efficient and cost-effective manner.

Future research should explore the application of this method to larger datasets, investigate ways to further refine the self-curation process, and examine the potential for combining instruction backtranslation with other techniques, such as self-alignment and data quality optimization. By continuing to develop and improve methods like instruction backtranslation, researchers can work towards creating more capable and versatile language models that can better understand and follow a wide range of instructions.


# Systematic Evaluation of Instruction-Tuned Large Language Models on Open Datasets

This <mark style="color:blue;">**October 2023**</mark> paper explores the impact of instruction-tuning on large language models using various open datasets.&#x20;

The authors aim to systematically evaluate the performance of these models across a range of tasks and provide a comprehensive comparison with state-of-the-art proprietary models like ChatGPT and GPT-4.

The researchers trained a series of instruction-tuned models, ranging from 6.7B to 65B parameters, on 12 different instruction datasets.&#x20;

We would argue that this paper does not take into account significant progress made on highly curated datasets, as referred to in the [<mark style="color:blue;">**Alpagasus paper**</mark>](/data/datasets/instruction-fine-tuning-alpagasus) and the [<mark style="color:blue;">**LIMA paper**</mark>](/data/datasets/less-is-more-for-alignment).

{% embed url="<https://arxiv.org/abs/2306.04751>" %}

### <mark style="color:purple;">Key findings</mark>

1. Different instruction-tuning datasets can enhance specific skills, but <mark style="color:yellow;">no single dataset or combination provides the best performance across all evaluations</mark>.
2. Larger or pretrained-for-longer base models consistently outperform smaller ones after instruction tuning.
3. Even the largest model (65B) finetuned on a mix of instruction datasets fails to outperform ChatGPT, although it significantly outperforms similar smaller models.
4. Model and human preference-based evaluations do not always reflect differences in model capabilities exposed by benchmark-based evaluations, highlighting the need for comprehensive evaluation.

### <mark style="color:purple;">Instruction tuning datasets</mark>

Below are some of the datasets tested.

To describe the table below:

* <mark style="color:blue;">**Dataset:**</mark> Dataset source
* <mark style="color:blue;">**Sourced from:**</mark> Where the data came from (human only or combination of model plus human)
* <mark style="color:blue;">**Instances:**</mark> The number of separate instructions in the dataset
* <mark style="color:blue;">**N rounds:**</mark> The average number of conversation turns in each dataset. A conversation turn consists of a user prompt and the corresponding assistant response.
* <mark style="color:blue;">**L prompt:**</mark> The average length (number of tokens) of the prompts in each dataset.
* <mark style="color:blue;">**L completion:**</mark> The average length (number of tokens) of the completions or responses generated by the assistant in each dataset.

These definitions provide a clear understanding of the terms used in the table and their relevance to the instruction datasets investigated in the research paper.

<table><thead><tr><th width="159">Datasets</th><th width="171">Sourced from</th><th width="108">Instances</th><th width="85">rounds</th><th width="93">prompt</th><th>completion</th></tr></thead><tbody><tr><td>CoT</td><td>NLP datasets + Human-written CoTs</td><td>100,000</td><td>1.0</td><td>266.0</td><td>53.2</td></tr><tr><td>FlanV2</td><td>NLP datasets + Human-written Instructions</td><td>100,000</td><td>1.0</td><td>355.7</td><td>31.2</td></tr><tr><td>Dolly</td><td>Human-written from scratch</td><td>15,011</td><td>1.0</td><td>118.1</td><td>91.3</td></tr><tr><td>OpenAssistant</td><td>Human-written from scratch</td><td>34,795</td><td>1.6</td><td>34.8</td><td>212.5</td></tr><tr><td>Self-instruct</td><td>Generated w/ vanilla GPT3</td><td>82,439</td><td>1.0</td><td>41.5</td><td>29.3</td></tr><tr><td>Unnatural Instructions</td><td>Generated w/ Davinci-002</td><td>68,478</td><td>1.0</td><td>107.8</td><td>23.6</td></tr><tr><td>Alpaca</td><td>Generated w/ Davinci-003</td><td>52,002</td><td>1.0</td><td>27.8</td><td>64.6</td></tr><tr><td>Code-Alpaca</td><td>Generated w/ Davinci-003</td><td>20,022</td><td>1.0</td><td>35.6</td><td>67.8</td></tr><tr><td>GPT4-Alpaca</td><td>Generated w/ Davinci-003 + GPT4</td><td>52,002</td><td>1.0</td><td>28.0</td><td>161.8</td></tr><tr><td>Baize</td><td>Generated w/ ChatGPT</td><td>210,311</td><td>3.1</td><td>17.6</td><td>52.8</td></tr><tr><td>ShareGPT3</td><td>User prompts + outputs from various models</td><td>168,864</td><td>3.2</td><td>71.0</td><td>357.8</td></tr></tbody></table>

### <mark style="color:purple;">Instruction Tuning Format</mark>

Instruction tuning is a technique used to finetune pretrained language models to better understand and respond to human requests expressed in natural language.&#x20;

An instruction tuning dataset consists of a collection of input-output pairs, where the input is a user prompt or instruction, and the output is the desired response or completion. The goal is to train the model to generate appropriate responses given the input prompts.

The format typically follows a chat-style schema, where the interactions between the user and the language model (assistant) are *<mark style="color:yellow;">**encoded together**</mark>* using special tokens.

The format consists of:

1. <mark style="color:purple;"><|user|></mark> token placed before user utterances
2. <mark style="color:purple;"><|assistant|></mark> token placed before target assistant responses
3. <mark style="color:purple;">\</s></mark> end-of-text marker placed at the end of each assistant output

To prepare the data for training, *<mark style="color:yellow;">**special tokens**</mark>* are used to demarcate the boundaries between the user prompts and the assistant responses.

In this case, <mark style="color:purple;">**<|user|>**</mark> is used to indicate the start of a user prompt, and <mark style="color:purple;">**<|assistant|>**</mark> is used to indicate the start of the assistant's response. &#x20;

The entire sequence, including the user prompts and assistant responses, is encoded together.

During training, the <mark style="color:yellow;">**loss is computed only on the tokens after**</mark>**&#x20;**<mark style="color:purple;">**<|assistant|>**</mark>**&#x20;**<mark style="color:yellow;">**and before the next**</mark>**&#x20;**<mark style="color:purple;">**<|user|>**</mark>**&#x20;**<mark style="color:yellow;">**token.**</mark>&#x20;

This is achieved using *<mark style="color:yellow;">**teacher-forcing with loss masking,**</mark>* where the tokens belonging to the input sequence(s) are masked.

Here's an example of the format in a code block:

{% code overflow="wrap" %}

```bash
<|user|> Explain the fault-tolerance of the reaction control system on the Space Shuttle.
<|assistant|> The reaction control system (RCS) on the Space Shuttle was designed to be fault-tolerant, meaning it was able to continue functioning even if one or more of its components failed. The RCS consisted of two sets of ... </s>

<|user|> Did the RCS have any on-orbit failures?
<|assistant|> There were several instances where the reaction control system (RCS) on the Space Shuttle experienced failures or malfunctions during on-orbit missions. These ... </s>
```

{% endcode %}

In this example, the user prompts are marked with <mark style="color:purple;">**<|user|>**</mark> tokens, and the assistant responses are marked with <mark style="color:purple;">**<|assistant|>**</mark> tokens. The <mark style="color:purple;">**\</s>**</mark> token is used to indicate the end of the assistant's output for each round.

During training, we compute loss only on tokens after <mark style="color:purple;">**`<|assistant|>`**</mark> and before the next <mark style="color:purple;">**`<|user|>`**</mark> token.&#x20;

More formally, we consider an instruction dataset as consisting of $$N$$ tuples, each with $$i$$ *<mark style="color:yellow;">**turns of conversation:**</mark>*

$$
{(x^1\_j, y^1\_j, x^2\_j, y^2\_j, \dots, x^i\_j, y^i\_j)}^N\_{j=1}$
$$

&#x20;where $$x\_i$$ is a *<mark style="color:yellow;">**user prompt**</mark>* and $$y\_i$$ the *<mark style="color:yellow;">**desired output.**</mark>*&#x20;

For most instances, $$i = 1$$ and we train the model to output $$y\_j$$ given $$x\_j$$ -  meaning the model is trained to generate the output $$y$$ given the input $$x$$.

However, in the case of conversation datasets, there can be multiple turns to train the model to predict $$y^i\_j$$ given some conversation history $$x^1\_j, y^1\_j, x^2\_j, \dots, x^i\_j$$.&#x20;

Given $$X$$ as the tokens belonging to the input, and $$Y$$ as the target tokens, the *<mark style="color:yellow;">**loss function**</mark>* is:

$$
L = -\sum\_j \log p\_\theta(t\_j \mid t\_{\<j}) \times \begin{cases} 1 & \text{if } t\_j \in Y \ 0 & \text{otherwise} \end{cases}
$$

where $$tⱼ$$ is the $$jth$$ input token (belonging to input $$X$$ or target $$Y$$.

The loss function used in this training process is the <mark style="color:yellow;">**negative log-likelihood loss**</mark>. This ensures that the *<mark style="color:yellow;">**loss is computed only on the tokens that are part of the assistant's response**</mark>*.

This is achieved by using <mark style="color:blue;">**teacher forcing with loss masking**</mark>.&#x20;

Teacher forcing means that the model is provided with the *<mark style="color:yellow;">**ground truth output at each step during training,**</mark>* and it learns to predict the next token based on the previous tokens.  Loss masking ensures that the loss is computed only on the relevant tokens (i.e., the assistant's response) and not on the input tokens.

### <mark style="color:purple;">The training process</mark>

Now, let's consider how this training process interacts with the Transformer architecture.&#x20;

The Transformer architecture consists of an encoder and a decoder, but in this case, a decoder-only model is used.

During training, the input sequence (user prompts and conversation history) is passed through the Transformer layers.&#x20;

The <mark style="color:blue;">**self-attention mechanism**</mark> in the Transformer allows the model to attend to different parts of the input sequence and capture the relevant information for generating the output.

At each <mark style="color:blue;">**decoding step**</mark>, the model takes the previously generated tokens as input and attends to the entire input sequence to predict the next token.  The loss is computed based on the predicted probability distribution over the vocabulary and the ground truth token at each step.

The Transformer's <mark style="color:blue;">**multi-head attention mechanism**</mark> enables the model to learn different aspects of the input-output relationship, allowing it to capture the nuances of the task. The feedforward layers in the Transformer help in processing and transforming the learned representations.

Through the <mark style="color:blue;">**process of fine-tuning**</mark>, the pretrained LLM *<mark style="color:yellow;">**adapts its parameters**</mark>* to the specific task of instruction following. The model learns to understand the patterns and relationships between the user prompts and the desired responses, enabling it to generate appropriate outputs for new, unseen prompts.

By training on a diverse set of instruction-following examples, the model becomes more versatile and can handle a wide range of tasks and conversations. The quality and diversity of the training dataset play a crucial role in the model's performance and generalization ability.

### <mark style="color:purple;">Decoder Only?</mark>

In the context of language modelling and sequence-to-sequence tasks, a decoder-only model refers to a specific architecture where only the decoder component of the Transformer is used, without an explicit encoder.

In a typical sequence-to-sequence model, the Transformer architecture consists of two main components: the encoder and the decoder. The encoder takes the input sequence and generates a set of hidden representations that capture the relevant information from the input. The decoder then takes these hidden representations and generates the output sequence step by step.

However, in a <mark style="color:blue;">**decoder-only model**</mark>, the entire input sequence (both the user prompts and the conversation history) is concatenated and fed directly into the decoder. The decoder attends to the input sequence and generates the output sequence in an autoregressive manner, meaning it predicts the next token based on the previously generated tokens.

The *<mark style="color:yellow;">**main difference between a decoder-only model and a standard encoder-decoder model**</mark>* is that the decoder-only model does not have a separate encoder to process the input sequence. Instead, the decoder itself attends to the input sequence and learns to generate the output based on the entire context.

### <mark style="color:purple;">Which is the best dataset?</mark>

The analysis of the instruction tuning datasets and base models reveals several key findings:

#### <mark style="color:green;">No single best dataset</mark>

There is no single instruction tuning dataset that performs best across all tasks.&#x20;

Different datasets enable different capabilities in the model, with notable examples being CoT for mathematical reasoning in GSM and Code-Alpaca for Codex-Eval.

#### <mark style="color:green;">Combining datasets is beneficial</mark>

Models trained on combined datasets generally achieve the best overall performance on benchmark tasks. While they may not be the best for individual tasks, they have the highest average performance across all tasks.

#### <mark style="color:green;">Base model quality is crucial</mark>

The choice of the base model significantly impacts downstream performance.&#x20;

LLAMA models outperform OPT and Pythia models of comparable size when trained on the same data mixture, likely due to LLAMA being pretrained on more tokens. &#x20;

The addition of LLAMA-2 further confirms this finding, showing that improvements can come from upgrading the base model alone.  This is no surprise.

### <mark style="color:purple;">Conclusion</mark>

In conclusion, this research provides a comprehensive evaluation of various publicly available resources for instruction tuning and compares their performance to state-of-the-art proprietary models like ChatGPT and GPT-4.&#x20;

The findings emphasise the importance of using strong base models, combining diverse datasets, and conducting thorough evaluations across a wide range of tasks and metrics.

The study highlights that no single instruction tuning dataset excels across all tasks, but rather different datasets enable different capabilities in the model. &#x20;

Combining datasets generally results in the best overall performance on benchmark tasks, although it may lead to slight performance drops compared to the best performance on specific tasks.

The quality of the base model plays a crucial role in downstream performance, with LLAMA models outperforming other models of comparable size when trained on the same data mixture.

However, despite the progress made, the strongest open models still fall short of matching the performance of proprietary models like ChatGPT and GPT-4. This gap underscores the need for continued development of robust base models and more diverse, comprehensive datasets.


# Instruction Tuning

Inspired by the Self-Instruct Paper

Instruction tuning is a specialised form of fine-tuning in which a <mark style="color:yellow;">model is trained using pairs of input-output instructions</mark>, enabling it to learn specific tasks guided by these instructions.

Instruction tuning strategies are techniques used to refine a language model's understanding and response to instructions. These *<mark style="color:yellow;">**strategies differ from pre-training**</mark>*, focusing on efficiency and targeted improvement.&#x20;

{% embed url="<https://arxiv.org/abs/2308.10792>" %}
Instruction Tuning
{% endembed %}

### <mark style="color:purple;">Instruction Tuning</mark>

Below are concise explanations of these strategies:

<mark style="color:green;">**Balancing Data Distribution**</mark>

This involves ensuring a proportional representation of tasks during instruction tuning to prevent any single dataset from dominating.&#x20;

Techniques include examples-proportional mixing, where instances are equally sampled from all datasets, and imposing a maximum cap on the number of examples per dataset to prevent data imbalance.

<mark style="color:green;">**Combining Instruction Tuning with Pre-Training**</mark>

This method enhances tuning effectiveness by mixing pre-training data (plain texts) with instruction-tuned data (formatted datasets), serving as regularization to prevent overfitting.&#x20;

This approach can either integrate instruction data during pre-training or combine both phases into one, using multi-task learning to benefit from both pre-training and instruction tuning simultaneously.

<mark style="color:green;">**Multi-stage Instruction Tuning**</mark>

A phased approach where the model is initially fine-tuned on task-formatted instructions (usually more abundant) and then on daily chat instructions.&#x20;

To mitigate the potential loss of previously learned information (capacity forgetting), task-formatted instructions are reintroduced in later stages. This tiered tuning can progressively introduce more complex and difficult tasks to incrementally challenge and improve the model.

<mark style="color:green;">**Data augmentation**</mark>

Augmenting the data such as by inverting inputs and outputs (e.g., turning a question answering task into a question generation task) is beneficial.

Neural language models can be enhanced to follow instructions on new, unseen tasks through a process called <mark style="color:yellow;">instruction tuning.</mark>&#x20;

### <mark style="color:purple;">Where did the strategy come from?</mark>

The concept of instruction tuning was first introduced in a widely cited paper from Google called  "Fine tuned language models are zero shot-learners" .

The paper focuses on enhancing the ability of neural language models to perform tasks they haven't been explicitly trained on, <mark style="color:yellow;">known as zero-shot learning.</mark>

This is where the term 'instruction tuning' was first coined, which has become the foundation for creating datasets for fine tuning foundation language models.

The process has been further refined and developed.  The paper below provides a comprehensive review of the history of the process:

### <mark style="color:purple;">Instruction Tuning</mark>

<mark style="color:green;">**Task Datasets**</mark>

These datasets are analogous to comprehensive workbooks comprising a variety of tasks—such as text summarisation, classification, and translation—each prefaced with natural language instructions.&#x20;

The role of these instructions is to clearly communicate the task's objective to the model, equipping it to perform the required function.  This method closely mimics a supervised learning environment where the presence of instructions is integral to guiding the model’s understanding and response to a task.

<mark style="color:green;">**Daily Chat Data**</mark>

To mimic real-world interaction and capture the diversity of human communication, training data is also sourced from everyday dialogues.&#x20;

These include a range of user queries that language models encounter when interacting with people in different contexts.&#x20;

This not only facilitates the instruction-following capability of the models but also aids in the refinement of their responses to align with actual human inquiries and needs.&#x20;

Additionally, this dataset includes human-generated instructions for a variety of real-life scenarios and the corresponding responses to construct a realistic dialogue training environment.

<mark style="color:green;">**Synthetic Data**</mark>

Recognising the limitations of solely relying on human-generated data, synthetic approaches are employed to augment training datasets.&#x20;

Language models are prompted to generate new task instructions and associated input-output pairs based on existing instances.&#x20;

This semi-automated process allows for the expansion of training materials without the excessive demand for human annotation, promoting a cost-effective method of enhancing the model's learning and generative capabilities.

These diverse data sources collectively contribute to the robustness of language models, ensuring that they are well-versed in both structured task-oriented interactions and the flexible, unpredictable nature of human dialogue.

Instruction tuning is a method of training large language models (LLMs) that involves <mark style="color:yellow;">augmenting traditional input-output data with explicit instructions</mark>, enhancing the model's ability to generalize to new tasks.&#x20;

<mark style="color:purple;">Below are the primary types of instruction tuning and the background research:</mark>

#### <mark style="color:green;">**Natural Instructions**</mark>

* Developed by Mishra et al. (2022), this dataset comprises 193,000 instruction-output examples derived from 61 English NLP tasks.
* The uniqueness lies in its structured approach, where instructions from each dataset are aligned to a common schema, including definitions, things to avoid, and examples.

#### <mark style="color:green;">**Super-Natural Instructions (Natural Instructions v2)**</mark>

* Created by Wang et al. (2022), it is an extension of the Natural Instructions dataset.
* It includes 5 million examples from 76 tasks in 55 languages, with instructions simplified to include task definitions and positive and negative examples with explanations.

#### <mark style="color:green;">**Unnatural Instructions**</mark>

* Introduced by Honovich et al. (2023), this dataset contains 240,000 examples generated by prompting InstructGPT (text-davinci-002) with Super-Natural Instructions examples.
* It covers a broader range of tasks than its predecessors and includes creative tasks beyond classical NLP challenges.

#### <mark style="color:green;">**Self-Instruct**</mark>

* Similar to Unnatural Instructions, this dataset by Wang et al. (2023) consists of 82,000 examples generated using InstructGPT.
* It decouples example generation into three steps: generating the instruction, then the input, and finally the output, aiming to reduce bias in classification tasks.

These datasets and approaches to instruction tuning highlight the evolving landscape of LLM training.&#x20;

By integrating explicit instructions into training data, instruction tuning enables models to understand and perform tasks more effectively, bridging the gap between machine understanding and human-like comprehension and reasoning.


# Instruction Fine Tuning - Alpagasus

"ALPAGASUS: Data-Driven Data Selection for Instruction Fine-Tuning"

Neural language models are improved through instruction-finetuning (IFT) - where they are trained on *<mark style="color:yellow;">**specific sets of instructions and corresponding responses.**</mark>*&#x20;

This paper highlights the importance of a curated and filtered dataset to improve the results from Instruction Fine-Tuning (IFT).

{% embed url="<https://arxiv.org/abs/2307.08701>" %}
AlpaGasus
{% endembed %}

The training of Meta's open sourced model LLama by Stanford University using a datasets of 52,000 instructions created a model called "Alpaca" (An Alpaca is a type of LLama)

The team that wrote the AlpaGasus paper found that the training data used by Stanford to create Alpaca contained many low-quality entries.&#x20;

They argued these entries can include incorrect or irrelevant responses, which *<mark style="color:yellow;">**negatively impact the finetuning process**</mark>*.

To address this, the team curated the Stanford dataset to filter out low-quality data from the original 52,000 Alpaca dataset.&#x20;

As a result, AlpaGasus model fine tunes Meta's LLama model on <mark style="color:green;">only 9,000 high-quality instances.</mark>&#x20;

Despite the smaller fine tuning dataset, the paper demonstrates that AlpaGasus significantly outperforms the original Alpaca model in various tests and human evaluations.

Alpagasus is evaluated on multiple test sets and compared to the original Alpaca and the results show it significantly outperforms Alpaca in terms of instruction-following capability.

Just as importantly, it also reduces training time substantially, making it more cost efficient.&#x20;

This proves the ongoing movement around data - to *<mark style="color:yellow;">**prioritise data quality over quantity.**</mark>*

### <mark style="color:purple;">Key Takeaways for Data Curation</mark>

<mark style="color:blue;">**Threshold-based Filtering**</mark>

Consider implementing a threshold-based filtering approach.  Set a threshold value for a scoring metric that reflects the quality of data. Data points with scores equal to or higher than the threshold are retained, while those below it are filtered out. This approach allows you to select data that meets a certain quality standard.

<mark style="color:blue;">**Comparative Filtering**</mark>

Explore the possibility of comparative filtering where you have multiple datasets or versions of a dataset. In the paper, they compared models trained on different subsets of data to assess the impact of data quality and quantity. You can create variations of your dataset, apply different filtering criteria, and then compare the performance of models trained on these subsets. This can help identify the best-performing subset of data.

<mark style="color:blue;">**Human vs. Machine Filtering**</mark>

Consider filtering datasets based on whether the data is human-generated or machine-generated. In the paper, they applied their filtering method to both machine-generated and human-written datasets. You can use this approach to assess the impact of data quality on models when using different data sources. It may be beneficial to have a separate filtering process for each type of data source.


# Less is More For Alignment

Co-authored by researchers from Meta, Carnegie Mellon University, University of Southern California, and Tel Aviv University

This <mark style="color:blue;">May 2023</mark> paper made waves in the AI community by challenging the prevailing notion that extensive fine-tuning is necessary for refining large language models (LLMs) and proposes a more efficient approach that leverages smaller, high-quality datasets.

The paper introduces LIMA, a 65-billion parameter LLaMa language model, which diverges from conventional training approaches by being fine-tuned with a standard supervised loss on only 1,000 carefully curated prompts and responses.&#x20;

Remarkably, LIMA demonstrates strong performance, learning to follow specific response formats from a minimal number of examples and generalising well to unseen tasks, ranging from planning trip itineraries to speculating about alternate histories.

In a controlled human study, LIMA's responses were either equivalent to or strictly preferred over those from GPT-4 in 43% of cases, a statistic that rises to 58% when compared against Bard and 65% versus DaVinci 003, models trained with extensive reinforcement learning and human feedback.&#x20;

These findings suggest that *<mark style="color:yellow;">**the bulk of knowledge in large language models is acquired during the pre training phase,**</mark>* and that only a limited amount of instruction tuning data is necessary for models to produce high-quality output. This challenges the prevailing notion that significant compute and specialised data are essential for achieving high performance, opening new avenues for efficient and effective language model training.

#### <mark style="color:green;">The Superficial Alignment Hypothesis</mark>

At the core of the paper is the "Superficial Alignment Hypothesis," which suggests that the majority of a language model's capabilities are acquired during the pretraining phase.&#x20;

According to this hypothesis, the primary role of alignment is to *<mark style="color:yellow;">**teach the model the appropriate format for user interaction.**</mark>* This idea shifts the focus from extensive fine-tuning to a more targeted approach that emphasizes dataset quality over quantity.

{% embed url="<https://arxiv.org/abs/2305.11206>" %}
"Less is More For Alignment": Challenging Conventions in Language Model Fine-Tuning
{% endembed %}

### <mark style="color:purple;">Hyperparameters Used</mark>

The training of LIMA incorporated a mix of standard and innovative hyperparameters, with specific attention to preventing overfitting and ensuring the model's ability to distinguish between different speakers. <mark style="color:yellow;">The manual selection of checkpoints highlights a focus on qualitative evaluation</mark>, which is especially important for models aimed at generating high-quality, contextually appropriate responses.

<mark style="color:green;">**Use of Special Token**</mark>

A unique end-of-turn (EOT) token was introduced to differentiate between user and assistant utterances. This choice avoids confusion with the existing end-of-sentence (EOS) token, which the pretrained model might associate with different meanings.

<mark style="color:green;">**Standard Fine-tuning Parameters**</mark>

The model was fine-tuned for 15 epochs using the AdamW optimizer. This optimizer is known for its effectiveness in large model training due to its handling of weight decay. The specific values for the AdamW hyperparameters were β1 = 0.9 and β2 = 0.95, with a weight decay of 0.1.

<mark style="color:green;">**Learning Rate Strategy**</mark>

The initial learning rate was set at 1e-5, with a linear decay to 1e-6 by the end of training. This gradual reduction is a common approach to stabilize and refine the learning process over time.

<mark style="color:green;">**Batch Size and Token Limit**</mark>

The batch size was set to 32 examples (or 64 for smaller models). Texts longer than 2048 tokens were trimmed, ensuring manageable training sizes and focusing the model's attention on the most relevant parts of the data.

<mark style="color:green;">**Residual Dropout**</mark>

A notable deviation from standard practice was the use of residual dropout. The dropout rate started at 0% at the bottom layer and increased linearly to 30% at the last layer (20% for smaller models). This approach, inspired by Ouyang et al. \[2022], suggests a focus on preventing overfitting in deeper layers of the model.

<mark style="color:green;">**Manual Selection of Checkpoints**</mark>

Interestingly, the team did not rely on perplexity as an indicator of generation quality. Instead, they manually selected checkpoints between the 5th and the 10th epochs using a held-out 50-example development set. This indicates a more hands-on approach to model tuning, prioritizing qualitative assessments over traditional metrics.

## <mark style="color:purple;">**Training Data**</mark>

The study demonstrated that LIMA, despite being trained on a limited yet diverse and high-quality dataset, was capable of robust performance, even outperforming models trained on much larger datasets in some cases.&#x20;

This underscores the importance of pretraining and the potential efficiency gains from fine-tuning with carefully selected data.

<mark style="color:green;">**Source and Composition**</mark>

The training data comprised *1,000 examples that mimic real user prompts and high-quality responses.* This set includes 750 top questions and answers from community forums like Stack Exchange (in both STEM and other categories) and wikiHow. These were selected for their quality and diversity. The average input and output lengths varied across sources, indicating a mix of brief and detailed interactions.

<mark style="color:green;">**Manual Contributions**</mark>

Additionally, the paper authors manually wrote 250 examples to optimise for task diversity and uniform response style. This was done to align with the spirit of an AI assistant, ensuring that the model could adapt to a variety of tasks while maintaining a consistent approach in its responses.

<mark style="color:green;">**Diverse Data Sources**</mark>

The training data encompassed a wide range of sources, including technical forums (Stack Exchange), practical guides (wikiHow), creative writing prompts (Pushshift r/WritingPrompts), and instructional examples (Natural Instructions). This diversity was aimed at exposing the model to a broad spectrum of query types and response styles.

<mark style="color:green;">**Controlled Data Size**</mark>

The total amount of training data was about 750,000 tokens, split over exactly 1,000 sequences. This small dataset size, especially compared to the extensive datasets typically used for language model training, was a deliberate choice to test the efficacy of fine-tuning on a highly curated dataset.

<mark style="color:green;">**Emphasis on Quality and Style**</mark>

The focus was on the quality and style of the data rather than its quantity. The training set aimed to teach LIMA how to interact with users effectively by following a specific format, rather than inundating it with an extensive amount of varied data.

The training set used in the study for LIMA, a 65B-parameter LLaMa model, was curated to test the hypothesis that alignment can be simplified by teaching the model to interact with users in a specific style, leveraging the knowledge and capabilities acquired during pretraining. Here's a detailed explanation of the training set:

### <mark style="color:purple;">Findings</mark>

<mark style="color:green;">**Dataset Quality Over Quantity**</mark><mark style="color:green;">:</mark> The study prioritises the quality of the dataset for fine-tuning, utilizing only 1,000 high-quality examples to achieve significant performance improvements.

<mark style="color:green;">**Dataset Composition and Size**</mark><mark style="color:green;">:</mark> Emphasizes the creation of a small yet diverse training dataset, including sources like Stack Exchange and WikiHow, to promote task diversity and a uniform response style.

<mark style="color:green;">**Diversity and Labourious Dataset Creation**</mark><mark style="color:green;">:</mark> Highlights the intensive effort required to create a diverse and high-quality dataset, underlining the importance of manual curation.

<mark style="color:green;">**Dataset Creation Challenges**</mark><mark style="color:green;">:</mark> Explores the challenges in effectively creating and refining datasets for fine-tuning, focusing on the balance between size and quality.

<mark style="color:green;">**Comparative Analysis with Other Models**</mark><mark style="color:green;">:</mark> LIMA is benchmarked against state-of-the-art models, showing comparable or superior performance in some instances.

<mark style="color:green;">**Performance on Out-of-Distribution Samples**</mark><mark style="color:green;">:</mark> LIMA demonstrates strong generalisation on novel tasks but shows a tendency for unsafe responses, highlighting the importance of diverse input for quality output.

<mark style="color:green;">**Single vs. Multi-Turn Dialogue Testing**</mark><mark style="color:green;">:</mark> Showcases LIMA's adaptability in various conversational contexts, from single-turn responses to coherent multi-turn dialogues.

<mark style="color:green;">**Annotation and Labeling Considerations**</mark><mark style="color:green;">:</mark> Discusses the evolving role of annotation and labeling, particularly in ensuring data quality and consistency.

<mark style="color:green;">**Factors Influencing Input Diversity**</mark><mark style="color:green;">:</mark> Considers multiple factors that contribute to dataset diversity, such as formatting, concepts, and sentiment.

<mark style="color:green;">**Measuring Input Diversity and Output Quality**</mark><mark style="color:green;">:</mark> Raises the challenge of quantitatively assessing critical factors like input diversity and output quality in datasets.

<mark style="color:green;">**Application-Specific Fine-Tuning**</mark><mark style="color:green;">:</mark> Considers the benefits of fine-tuning models for specific applications, suggesting a tailored approach based on domain requirements.

<mark style="color:green;">**Alignment Data and Superficial Alignment Hypothesis**</mark><mark style="color:green;">:</mark> Tests the hypothesis that alignment mainly involves teaching the model user interaction styles, leveraging pretraining knowledge.

<mark style="color:green;">**Prompt Creation Process**</mark><mark style="color:green;">:</mark> Highlights the meticulous process of creating prompts and responses to ensure uniformity and relevance.

<mark style="color:green;">**Interactivity and Intelligence**</mark><mark style="color:green;">:</mark> Suggests that model intelligence appears to enhance with longer user interactions and self-referencing.

### <mark style="color:purple;">Conclusion</mark>

&#x20;"Less is More For Alignment" presents a compelling case for efficient fine-tuning of LLMs using smaller, high-quality datasets.&#x20;

The paper challenges conventional wisdom and opens up new avenues for exploration in the field. As the landscape of LLMs continues to evolve rapidly, this research contributes to the ongoing discourse on model refinement and alignment.

The paper highlights the importance of dataset quality, diversity, and the need for comprehensive evaluation metrics. It also underscores the potential for leveraging LLMs themselves in model assessment and the evolving role of annotation and labeling in dataset curation.

As the AI community continues to push the boundaries of what's possible with LLMs, papers like "Less is More For Alignment" inspire researchers and practitioners to rethink assumptions and explore innovative approaches. The insights gained from this research will undoubtedly shape the future of language model fine-tuning and contribute to the advancement of this transformative technology.


# Enhanced Supervised Fine Tuning

How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition

This <mark style="color:blue;">**January 2024**</mark> paper investigates the interplay of data composition during supervised fine-tuning (SFT) of large language models (LLMs) to *<mark style="color:yellow;">**enhance their abilities in mathematical reasoning, code generation, and general human-aligning tasks.**</mark>*&#x20;

{% embed url="<https://arxiv.org/abs/2310.05492>" %}
How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition
{% endembed %}

The authors explore four key research questions related to model performance, data amount, composition ratio, model size, and Supervised Fine Tuning (SFT) strategies.

### <mark style="color:purple;">Impact of different Supervised Fine Tuning strategies</mark>

The authors explore four SFT strategies: multi-task learning, sequential training, mixed sequential training, and <mark style="color:yellow;">**Dual-stage Mixed Fine-tuning (DMT).**</mark>

* Multi-task learning leads to conflicts, while sequential training results in catastrophic forgetting.
* The proposed *<mark style="color:yellow;">**DMT strategy effectively alleviates both performance conflicts and catastrophic forgetting by balancing general and specialized abilities.**</mark>*

<figure><img src="/files/Ht0luzuuwYopBt4c6WHN" alt=""><figcaption><p>The illustration of four different training strategies in this paper</p></figcaption></figure>

### <mark style="color:purple;">Process for emulating their experiment</mark>

1. Collect the required datasets (GSM8K RFT, Code Alpaca, and ShareGPT) and evaluation benchmarks (GSM8Ktest set, HumanEval, and MT-Bench).
2. Prepare the pre-trained LLaMA models (7B, 13B, 33B).
3. For each research question:&#x20;

a. Create the necessary data subsets according to the experimental design.

&#x20;b. Fine-tune the LLaMA models using the specified training strategies and hyperparameters.

&#x20;c. Evaluate the fine-tuned models on the corresponding benchmarks.&#x20;

d. Analyze the results and compare the performance across different settings.

```python
# Pseudo-code for fine-tuning and evaluation

# Fine-tuning
for model_size in [7, 13, 33]:
    for data_subset in data_subsets:
        model = load_pretrained_model(f"llama-{model_size}b")
        fine_tuned_model = fine_tune(model, data_subset, epochs=3, learning_rate=2e-5, batch_size=16)
        save_model(fine_tuned_model)

# Evaluation
for model_size in [7, 13, 33]:
    for fine_tuned_model in fine_tuned_models:
        gsm8k_score = evaluate(fine_tuned_model, gsm8k_test_set)
        human_eval_score = evaluate(fine_tuned_model, human_eval)
        mt_bench_score = evaluate(fine_tuned_model, mt_bench)
        log_results(model_size, fine_tuned_model, gsm8k_score, human_eval_score, mt_bench_score)
```

### <mark style="color:purple;">Discussion</mark>

#### <mark style="color:green;">Visualisation of Different SFT Abilities</mark>

* The authors visualise the semantic representations of different SFT abilities using <mark style="color:blue;">**t-SNE**</mark>.
* They observe a collapse phenomenon in the semantic representations of both the original LLaMA-13b and LLaMA-13b with DMT.
* While there is some separation in the mathematical data representations, there is still overlap between code and general samples.

#### <mark style="color:green;">Ablation of the Specialised Domains in ShareGPT</mark>

The authors investigate the impact of removing code and math-related samples from the ShareGPT dataset on the performance gains observed in mixed data settings.&#x20;

They want to determine whether the performance improvements in low-resource scenarios are solely due to the presence of code and math samples in ShareGPT or if other factors contribute to these gains.

To do this, they use an open-set tagger (InsTag) to annotate samples in ShareGPT and remove instances containing the keywords "code" or "math" using regular expression matching. This process *<mark style="color:yellow;">**reduces the ShareGPT dataset from 86K to 63K samples**</mark>*.

They then conduct experiments - using different proportions of the modified ShareGPT dataset (without code and math) mixed with GSM8K and Code Alpaca datasets. &#x20;

The results show that removing code and math samples from ShareGPT not only mitigates performance conflicts among different abilities under high-resource conditions but also maintains stable gains in low-resource settings.

This finding suggests that the diversity and variability of the data in ShareGPT, rather than the specific code and math samples, contribute to the performance improvements in low-resource scenarios.&#x20;

The presence of code and math data within ShareGPT is not the key factor driving the performance gains identified in Section 3.3, which highlights the generalization of the conclusions.

#### <mark style="color:green;">Specialised Data Amount in Dual-stage Mixed Fine-tuning (DMT)</mark>

In this section, the authors explore how different values of k (<mark style="color:yellow;">the proportion of specialised data</mark>) influence model performance in the Dual-stage Mixed Fine-tuning (DMT) strategy.&#x20;

They adjust the value of k from 0 to 1 and observe the following:

1. When k increases from 0 to 1/256, the SFT models show significant improvements in both specialised ability and general human-aligning ability.
2. As k increases from 1/4 to 1, the model exhibits a decline in general ability, consistent with the findings that high-resource settings lead to conflicts.
3. When k increases from 1/256 to 1/4, there is a linear inverse trend between general ability and specialized ability, with an increase in general ability coinciding with a decrease in specialised ability.

These observations suggest that the value of k needs to be tuned based on specific requirements to achieve a balance between multiple abilities.&#x20;

### <mark style="color:purple;">Practical applications of the paper's findings</mark>

1. Improving the performance of LLMs in specialised domains by leveraging the DMT strategy and carefully tuning the amount of specialised data.
2. Developing more efficient and effective training strategies for LLMs to acquire multiple abilities while minimising performance conflicts and catastrophic forgetting.
3. Enhancing the adaptability of LLMs to low-resource settings by leveraging diverse and variable data sources.
4. Designing LLMs that can effectively balance and switch between general and specialised abilities based on the specific requirements of the task at hand.

### <mark style="color:purple;">Dual-stage Mixed Fine-tuning (DMT)</mark>&#x20;

Dual-stage Mixed Fine-tuning (DMT) is a training strategy proposed in the paper to *<mark style="color:yellow;">**address the challenges of ability conflicts during multi-task learning and catastrophic forgetting during sequential training.**</mark>*&#x20;

The key idea behind DMT is to <mark style="color:yellow;">**first learn a large amount of specialised data**</mark> and then add a small amount of specialised data to the general data during the final stage of fine-tuning to prevent forgetting.

#### <mark style="color:green;">**The DMT process consists of two stages**</mark>

<mark style="color:purple;">**Stage 1**</mark>

Fine-tune the pre-trained language model (LLaMA) on the specialised datasets (e.g., GSM8K RFT for math reasoning and Code Alpaca for code generation) using supervised fine-tuning (SFT). &#x20;

This stage is similar to the first stage of the mixed sequential training strategy.

<mark style="color:purple;">**Stage 2**</mark>

Fine-tune the model from Stage 1 using a mixed data source that combines the general data (e.g., ShareGPT) with varying proportions (k) of the specialized data (code and math).&#x20;

The values of k can be 1, 1/2, 1/4, 1/8, 1/16, or 1/32. Adding a small amount of specialized data in this stage helps the model recall the specialized abilities learned in Stage 1.

Here's a simplified code structure to emulate the DMT process:

```python
# Prepare the datasets
specialized_datasets = ["gsm8k_rft", "code_alpaca"]
general_dataset = "sharegpt"

# Stage 1: Specialized fine-tuning
for dataset in specialized_datasets:
    model = pretrained_llama_model()
    fine_tuned_model = fine_tune(model, dataset)
    save_model(fine_tuned_model, f"{dataset}_stage1")

# Stage 2: Mixed fine-tuning
k_values = [1, 1/2, 1/4, 1/8, 1/16, 1/32]
for dataset in specialized_datasets:
    for k in k_values:
        mixed_dataset = create_mixed_dataset(general_dataset, dataset, k)
        model = load_model(f"{dataset}_stage1")
        fine_tuned_model = fine_tune(model, mixed_dataset)
        save_model(fine_tuned_model, f"{dataset}_stage2_k{k}")

# Evaluation
for dataset in specialized_datasets:
    for k in k_values:
        model = load_model(f"{dataset}_stage2_k{k}")
        evaluate(model, dataset)
```

Note: The code above is a simplified representation of the process and would need to be adapted to work with the specific libraries and frameworks used for fine-tuning and evaluation.

By following this process, you can emulate the DMT strategy and investigate its effectiveness in mitigating catastrophic forgetting and achieving a balance between specialized and general abilities in the fine-tuned language models.


# Visualising Data using t-SNE

This highly cited <mark style="color:blue;">**2008**</mark> paper presented <mark style="color:blue;">**t-SNE (t-distributed Stochastic Neighbor Embedding)**</mark>, an advanced technique for *<mark style="color:yellow;">**visualising high-dimensional data by mapping it onto a two or three-dimensional space.**</mark>*&#x20;

This method is an evolution of the original <mark style="color:blue;">**Stochastic Neighbor Embedding (SNE)**</mark> developed by Hinton and Roweis in 2002. &#x20;

t-SNE modifies SNE to enhance the visualisation quality and ease of optimisation, addressing particularly the issue of crowding points in the centre of the map.&#x20;

This is especially crucial for data lying across multiple, related low-dimensional manifolds, common in datasets like images from various perspectives or text data.

{% embed url="<https://arxiv.org/abs/2108.01301>" %}
Visualizing Data using t-SNE
{% endembed %}

### <mark style="color:purple;">Key Contributions and Methodology</mark>

<mark style="color:green;">**Improved Optimisation:**</mark> t-SNE is easier to optimise compared to its predecessor SNE.

<mark style="color:green;">**Better Visualization:**</mark> The technique reduces crowding at the map's centre, a common issue in similar methods, which enhances the visualisation's readability and effectiveness.

<mark style="color:green;">**Adaptability:**</mark> t-SNE can visualise complex data structures from various domains, adapting to the intrinsic scales and densities of the data.

### <mark style="color:purple;">Technical Details</mark>

<mark style="color:green;">**Probability Distributions**</mark>

t-SNE starts by converting high-dimensional Euclidean distances between data points into conditional probabilities that express similarities.  These probabilities help in maintaining local structures of the data in the lower-dimensional space.

<mark style="color:green;">**Kullback-Leibler Divergence**</mark>

t-SNE minimises the sum of the Kullback-Leibler divergences between the joint probabilities of the high-dimensional and low-dimensional spaces, effectively keeping similar data points close in the map while allowing dissimilar points to be farther apart.

<mark style="color:green;">**Gradient Descent**</mark>

The method uses gradient descent to find the map that best represents the high-dimensional data's structure. The gradient terms are derived based on the difference in probabilities, with additional momentum and noise terms to optimize the embedding effectively.

<figure><img src="/files/vboPj8QWQKacjGVdayRl" alt=""><figcaption><p>Visualisations of 6,000 hand written digits from the MNIST dataset</p></figcaption></figure>

### <mark style="color:purple;">Practical Impact and Theoretical Significance</mark>

The practical impact of t-SNE is profound, as it significantly improves the visualisation of complex datasets with intricate internal structures.&#x20;

The theoretical implications include a better understanding of how *<mark style="color:yellow;">**dimensionality reduction can be effectively achieved by managing the trade-offs between local and global data structures.**</mark>* This makes t-SNE particularly useful for datasets where preserving both types of structures is crucial for meaningful analysis.

&#x20;t-SNE advanced the field of dimensionality reduction by introducing robust methods to handle the inherent complexities of visualizing high-dimensional data. These enhancements make it a preferred tool in many applications, ranging from bioinformatics to social network analysis.

Conclusion:

t-SNE is a powerful tool for visualizing high-dimensional data effectively, particularly useful in domains where the data's intrinsic structure is complex and multi-scaled. Despite its computational demands and sensitivity to parameter settings, t-SNE's ability to produce superior visualizations makes it a valuable method in the toolbox of machine learning practitioners and data scientists.

### <mark style="color:purple;">Conclusions and Future Work</mark>

The paper concluded that t-SNE is highly effective for visualising complex datasets by retaining local data structures while revealing global structures like clusters.&#x20;

The technique is computationally intensive, but methods like the landmark approach help in managing these demands.

### <mark style="color:purple;">Software Libraries Implementing t-SNE</mark>

t-SNE is widely implemented across several major machine learning and data analysis libraries, including:

<mark style="color:green;">**Scikit-learn (Python):**</mark> Provides a well-optimised implementation of t-SNE, commonly used in academia and industry for data visualisation tasks.

<mark style="color:green;">**R (Rtsne package):**</mark> Offers an implementation tailored for use within the R statistical computing environment.

<mark style="color:green;">**MATLAB:**</mark> Includes t-SNE functions in its Statistics and Machine Learning Toolbox, facilitating easy integration with other MATLAB functionalities.

### <mark style="color:purple;">Modern Applications of t-SNE</mark>

t-SNE is used across various fields to analyse and visualise high-dimensional data:

#### <mark style="color:green;">**Biomedical Data Visualisation**</mark>

For example, it is used in single-cell RNA sequencing data analysis to visualise the variation in gene expression levels across individual cells, helping identify different cell types based on their gene expression profiles.

#### <mark style="color:green;">**Financial Data Analysis**</mark>

Analysts use t-SNE to identify clusters of similar financial products or to analyze consumer behavior based on high-dimensional data.

#### <mark style="color:green;">**Image Data Exploration**</mark>

t-SNE helps in visualizing datasets of high-resolution images, grouping similar images together, which is useful in fields like digital pathology or retail catalogue management.

t-SNE continues to be a vital tool in machine learning and data science, with ongoing research aimed at improving its theoretical understanding and computational efficiency. Its ability to reveal intricate structures hidden within complex datasets makes it an indispensable tool for exploratory data analysis.


# UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

<mark style="color:blue;">**UMAP (Uniform Manifold Approximation and Projection)**</mark> is another popular technique for dimensionality reduction and visualisation, similar to <mark style="color:blue;">**t-SNE (t-Distributed Stochastic Neighbor Embedding)**</mark>.&#x20;

This highly cited <mark style="color:blue;">**September 2020**</mark> paper introduced UMAP (Uniform Manifold Approximation and Projection), a novel manifold learning technique for dimension reduction that builds upon theoretical foundations in Riemannian geometry and algebraic topology.&#x20;

{% embed url="<https://arxiv.org/abs/1802.03426>" %}
UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
{% endembed %}

### <mark style="color:purple;">Summary of UMAP Paper</mark>

#### <mark style="color:green;">**Introduction and Motivation**</mark>

UMAP is described as a scalable algorithm applicable to real-world data across various fields, with a particular emphasis on handling large datasets more efficiently than existing methods like t-SNE.&#x20;

UMAP is presented as a robust, general-purpose dimension reduction technique that not only excels in visualisation quality comparable to t-SNE but also offers superior runtime performance and flexibility in handling higher-dimensional embeddings.

#### <mark style="color:green;">**Theoretical Background**</mark>

The core of UMAP's methodology is grounded in a sophisticated mathematical framework involving Riemannian geometry and category theory, specifically leveraging the geometric realization of fuzzy simplicial sets.&#x20;

This approach allows UMAP to effectively maintain both local and global data structures, which is a significant advantage over methods that may only preserve one at the expense of the other.

#### <mark style="color:green;">**Algorithm Details**</mark>

UMAP’s algorithm is designed to approximate a high-dimensional manifold using a low-dimensional projection.&#x20;

It emphasizes the preservation of local structures through a process that initially estimates how data points are interconnected in high-dimensional space and then optimally projects these points onto a lower-dimensional space.

#### <mark style="color:green;">**Practical Implementation**</mark>

The paper discusses the implementation details of UMAP, including algorithmic steps and the impact of various hyperparameters.&#x20;

This section is crucial for practitioners who need to understand how to apply UMAP effectively to real datasets and how to tune it to achieve the best performance.

#### <mark style="color:green;">**Performance Evaluation**</mark>

UMAP is evaluated against other dimension reduction techniques, demonstrating its effectiveness through practical results on real-world datasets.&#x20;

These comparisons highlight UMAP's efficiency in scaling to large datasets and its ability to produce high-quality visualizations.

#### <mark style="color:green;">**Discussion of Limitations and Extensions**</mark>

The paper concludes with a discussion on the relative weaknesses of UMAP and scenarios where it might not be the optimal choice.

It also explores potential extensions of the algorithm, such as applications in semi-supervised learning, metric learning, and heterogeneous data embedding, suggesting avenues for future research and development.

**Contributions and Theoretical Justifications**

One of the primary contributions of this work is reframing dimension reduction and manifold learning problems within a new mathematical language, providing a fresh perspective that enhances theoretical understanding and practical application.

#### Conclusion

Overall, the UMAP paper presents a significant advancement in the field of dimension reduction, offering a method that balances the preservation of local and global structures with computational efficiency. Its robust theoretical underpinnings and practical effectiveness make it a valuable tool for data scientists and researchers working with complex high-dimensional data.

### <mark style="color:purple;">Difference between t-SNE and UMAP</mark>

Both are used to visualise complex high-dimensional data in a lower-dimensional space (typically two or three dimensions). However, they differ in several key aspects:

#### <mark style="color:green;">Theoretical Foundations</mark>

* **t-SNE:** Focuses on preserving the local structure of the data and tends to form tighter clusters than UMAP. It uses a probability distribution based on Gaussian distributions in the high-dimensional space and a Student-t distribution in the low-dimensional space to model the relationships between points.
* **UMAP:** Based on a mathematical framework from topology and algebraic geometry, UMAP uses concepts from Riemannian geometry and algebraic topology to project data. It approximates a high-dimensional manifold using a fuzzy topological structure and then uses graph layout algorithms to structure data in low dimensions.

#### <mark style="color:green;">Performance and Scalability</mark>

* **t-SNE:** Known for its ability to create visually appealing maps that reveal structure at many different scales. It is computationally intensive, especially as the size of the dataset grows, which can make it less scalable for very large datasets.
* **UMAP:** Often faster than t-SNE and scales better to larger datasets. UMAP's performance makes it suitable for larger datasets where t-SNE might struggle both in terms of computational resources and runtime.

#### <mark style="color:green;">Flexibility and Generalisation</mark>

* **t-SNE:** Has a primary focus on visualization and does not inherently support new data points (out-of-sample data) without retraining the whole model.
* **UMAP:** More flexible in general usage beyond visualization, such as in clustering, data preprocessing, and even semi-supervised learning tasks. UMAP can handle new (out-of-sample) data without needing to retrain the model, which is a significant advantage in dynamic datasets.

#### <mark style="color:green;">Visualization Quality</mark>

* **t-SNE:** Excellent at revealing local data structures and relationships in the dataset. It is particularly good at separating clusters clearly, which makes it highly effective for exploratory data analysis.
* **UMAP:** Also produces high-quality visualisations but tends to preserve more of the global structure compared to t-SNE. This can make UMAP more useful in cases where understanding the overall data structure is as important as understanding local groupings.

#### <mark style="color:green;">Parameters and Stability</mark>

* **t-SNE:** The results can be quite sensitive to the choice of perplexity and other hyperparameters. Different runs can sometimes yield different visualizations, which may require some tweaking to get consistent results.
* **UMAP:** Generally more robust to the choice of its parameters. While it still requires parameter tuning (like the number of neighbours), it typically provides more consistent results across different runs than t-SNE.

In summary, UMAP tends to be faster, more scalable, and versatile for a broader range of tasks beyond just visualisation, compared to t-SNE.&#x20;

However, t-SNE may still be preferred when the primary goal is to explore the local structure of the data with high granularity, especially in smaller datasets.


# Training and Evaluation Datasets

#### Random Splitting

* **Implementation:** <mark style="color:blue;">Divide the dataset into training and validation sets randomly</mark>. This method assumes that the data is uniformly distributed and that a random sample will be representative of the whole.
* **Benefits:** Ensures a mix of all types of data in both sets, preventing model overfitting to specific patterns only present in the training set.
* **Monitoring Overfitting:** By using a random validation set, you can monitor if the model performs significantly better on the training data compared to the validation data, indicating overfitting.

#### Time-based Splitting

* **Implementation:** In datasets where the temporal aspect is critical (e.g., news articles, financial data), the <mark style="color:blue;">data is split based on a certain time.</mark> For example, training on data from previous years and validating on the most recent year.
* **Avoiding Data Leakage:** This method is crucial for preventing data leakage, where the model inadvertently learns from future information not supposed to be available at the training time.
* **Real-world Performance:** Time-based splitting helps evaluate how the model will perform in real-world scenarios, dealing with recent or future data it hasn't been exposed to during training.

#### Stratified Splitting

* **Implementation:** Stratify the dataset based on categories or classes to ensure that each category is represented proportionally in both training and validation sets. This is particularly important in datasets with imbalanced classes.
* **Class-wise Performance:** It allows for a more detailed analysis of the model's performance across different categories, identifying if it struggles with particular types of data.
* **Bias Mitigation:** Helps in mitigating biases by ensuring that minority classes are adequately represented and evaluated.

#### Additional Techniques and Considerations:

**Cross-Validation**

* **Process:** Involves dividing the dataset into several subsets and rotating these subsets as training and validation sets. This method is especially useful for small datasets.
* **Comprehensive Evaluation:** Provides a thorough assessment of the model's performance across different subsets of data, offering a more robust estimate of its generalization ability.

**Leave-One-Out Strategy**

* **Concept:** A form of cross-validation where each data point is used as a single validation set, and the rest as training data. This is computationally intensive but can be insightful for small and critical datasets.

**Monitoring Techniques**

* **Learning Curves:** Plot learning curves that show the model's performance on the training and validation datasets over epochs. Divergence of these curves suggests overfitting.
* **Early Stopping:** Implement early stopping to halt the training when the model's performance on the validation set ceases to improve, preventing overfitting.

**Automated Splitting and Evaluation Tools**

* Utilize machine learning frameworks that offer automated data splitting and evaluation tools, providing insights into model performance and potential issues like data leakage or imbalance.

#### Conclusion

Each splitting strategy has its strengths and is suited to particular types of datasets and use cases. The key is to match the splitting strategy to the nature of your data and the specific requirements of your application. In your role, focusing on Next.js and LLMs, integrating these strategies effectively can significantly enhance the robustness and reliability of the models you develop. It's not just about how well the model learns the training data, but more importantly, how well it generalizes to new, unseen data.


# What is perplexity?

Perplexity is a commonly used evaluation metric in natural language processing (NLP) that measures *<mark style="color:yellow;">**how well a language model predicts a sample of text**</mark>*.&#x20;

Perplexity is a measurement of how well a probability distribution or a probability model predicts a sample.&#x20;

In the context of language modeling, perplexity measures how well a language model predicts the next word in a sequence based on the words that come before it.

Mathematically, perplexity is defined as the exponential of the average negative log-likelihood of a sequence of words. The formula for perplexity is:

Perplexity = $$exp(-1/N \* sum(log(P(w\_i|w\_1, w\_2, ..., w\_{i-1}))))$$

where:

* N is the total number of words in the sequence
* $$P(w\_i|w\_1, w\_2, ..., w\_{i-1})$$ is the probability of the word $$w\_i$$ given the preceding words $$w\_1, w\_2, ..., w\_{i-1}$$
* log is the natural logarithm

### <mark style="color:purple;">Intuitive Understanding</mark>

Perplexity can be thought of as a measure of how "surprised" or "confused" the language model is when predicting the next word.&#x20;

*<mark style="color:yellow;">**A lower perplexity indicates that the model is less surprised**</mark>* and can predict the next word more accurately, while a higher perplexity suggests that the model is more uncertain or confused.

For example, if a language model has a perplexity of 10 on a given text dataset, it means that, on average, the model is as confused as if it had to choose uniformly and independently from 10 possibilities for each word.

### <mark style="color:purple;">Technical Explanation</mark>

To calculate perplexity, you first need to *<mark style="color:yellow;">**compute the cross-entropy loss**</mark>* between the predicted word probabilities and the actual word probabilities. <mark style="color:blue;">**Cross-entropy loss**</mark> measures the *<mark style="color:yellow;">**difference between two probability distributions**</mark>*.

In the context of language modeling, the model predicts the probability distribution over the vocabulary for the next word, given the preceding words. The actual word distribution is represented as a one-hot vector, where the correct word has a probability of 1, and all other words have a probability of 0.

The cross-entropy loss for a single word is calculated as:

Loss = $$-log(P(w\_i|w\_1, w\_2, ..., w\_{i-1}))$$

To get the average cross-entropy loss for the entire sequence, you sum up the individual word losses and divide by the total number of words:

Average Loss = $$-1/N \* sum(log(P(w\_i|w\_1, w\_2, ..., w\_{i-1})))$$

Finally, perplexity is obtained by exponentiating the average cross-entropy loss:

Perplexity = $$exp(Average Loss)$$

The perplexity score is often used to compare different language models or to evaluate the improvement of a model during training.&#x20;

A lower perplexity indicates better language modeling performance.

It's important to note that while perplexity is a useful metric, it *<mark style="color:yellow;">**has some limitations.**</mark>*&#x20;

It doesn't directly measure the quality or coherence of the generated text, and it can be sensitive to the choice of vocabulary and the specifics of the training data.&#x20;

Therefore, *<mark style="color:yellow;">**perplexity should be used in conjunction with other evaluation metrics**</mark>* and human judgment to assess the overall performance of a language model.


# Foundation Models

Continuum AI Driven Modules

At the February 2024 World Government Summit in Dubai, NVIDIA's CEO Jensen Huang was pushing the concept of sovereign AI.

He stressed the importance of countries, including developing nations, to invest in AI infrastructure and generative AI that reflects their unique cultural and linguistic heritage.

While mildly self-serving - we believe he is correct. Every sovereign should be developing an AI infrastructure, deep learning capabilities and ultimately proprietary generative AI models.

It is not just about preserving culture and language - it is about building AI capability so it can be made available to the broader community.

We are looking for greater leadership from the Australian Government.

The CSIRO through Data61 and the Universities are making great strides, but the capital requirements to develop a sovereign AI platform are significant.

The Federal and State Government could supercharge our journey. Not through handouts, but through incentives and community engagement.

We are of course happy to help. :-).  Click on the link below for more information:

[<mark style="color:blue;">www.continuumlabs.ai</mark>](https://training.continuumlabs.ai/models/www.continuumlabs.ai)

### <mark style="color:purple;">Middle East</mark>

The Middle East are running hard at developing their AI capability. &#x20;

Abu Dhabi has already developed a high quality artificial intelligence model called "Falcon" which is  open source, so available for for research and commercial use.

The UAE have made some solid progress, having developed an early stage large language modell

Saudia Arabia are buying thousands of the latest H-100 GPU chips.

### <mark style="color:purple;">United Kingdom</mark>

The UK government has undertaken a range of policy initiatives in recent years to develop its 'strategic approach' to the use of AI.&#x20;

This includes the establishment of a Foundation Model Taskforce and funding of £100 million to deliver the government’s ambitions for UK capability in safe and reliable foundation models.

This initiative is in addition to investment of around £900 million for ‘a new ‘exascale’ supercomputer and a dedicated AI Research Resource to equip the UK with the processing power it needs to support the next generation of AI innovation.

### <mark style="color:purple;">Singapore</mark>

Singapore is seeking to advance its Artificial Intelligence (AI) capabilities through the National Multimodal Large Language Model Programme - a collaborative effort involving a range of Singaporean organisations.  &#x20;

Their goal is to make Singapore a global AI leader by 2030 - a global AI hub.

The key point is that a cornerstone of this initiative is the development of multimodal and localised large language models that *<mark style="color:yellow;">**reflect the diverse cultures and languages of Southeast Asia**</mark>*.

We expect there will be many more announcements from sovereigns developing their own generative AI platforms over the coming 12 months.

We would like to see Australia get involved.


# The leaderboard

### About

With the rapid release of numerous large language models (LLMs) and chatbots, often accompanied by bold claims regarding their performance, it can be challenging to discern genuine progress from the open-source community and identify the current state-of-the-art models.&#x20;

The Huggingface leaderboard addresses this need by providing a transparent and standardised evaluation of these models.

{% embed url="<https://huggingface.co/open-llm-leaderboard>" %}

They use the Eleuther AI Language Model Evaluation Harness, a unified framework for testing generative language models on various tasks, to evaluate the models on six key benchmarks.&#x20;

Detailed information about the evaluation tasks, results, and how to reproduce them is provided below.

### <mark style="color:purple;">Evaluation Tasks</mark>

The models are evaluated on the following six benchmarks:

1. <mark style="color:blue;">**IFEval**</mark>
   * **Description**: Tests the model’s ability to <mark style="color:yellow;">follow explicit instructions</mark>, focusing on formatting adherence.
   * **Shots**: 0-shot.
2. <mark style="color:blue;">**Big Bench Hard (BBH)**</mark>
   * **Description**: Evaluates models on 23 challenging tasks from the BigBench dataset.
   * **Shots**: 3-shot.
   * **Subtasks** include sports understanding, object tracking, logical deduction, and more.
3. <mark style="color:blue;">**MATH Level 5**</mark>
   * **Description**: Compiles high-school level competition problems requiring specific output formatting.
   * **Shots**: 4-shot.
4. <mark style="color:blue;">**Graduate-Level Google-Proof Q\&A Benchmark (GPQA)**</mark>
   * **Description**: Contains challenging knowledge questions crafted by PhD-level experts in various fields.
   * **Shots**: 0-shot.
5. <mark style="color:blue;">**Multistep Soft Reasoning (MuSR)**</mark>
   * **Description**: Consists of complex problems requiring reasoning and long-range context parsing.
   * **Subtasks**: 0-shot evaluation on murder mysteries, object placement, and team allocation.
6. <mark style="color:blue;">**Massive Multitask Language Understanding - Professional (MMLU-PRO)**</mark>
   * **Description**: A refined version of the MMLU dataset, featuring more challenging and noise-reduced questions.
   * **Shots**: 5-shot.

### <mark style="color:purple;">Model Comparison</mark>

Here is a detailed comparison of different language models across various evaluation categories:

<table data-view="cards"><thead><tr><th>Model Name</th><th>Average</th><th>Multi choice</th><th>Reasoning</th><th>Coding</th><th>Future Capabilities</th><th>Grade School Math</th><th>Math Problems</th></tr></thead><tbody><tr><td><strong>Claude 3.5 Sonnet</strong></td><td>88.38%</td><td>88.70%</td><td>89.00%</td><td>92.00%</td><td>93.10%</td><td>96.40%</td><td>71.10%</td></tr><tr><td><strong>Claude 3 Opus</strong></td><td>84.83%</td><td>86.80%</td><td>95.40%</td><td>84.90%</td><td>86.80%</td><td>95.00%</td><td>60.10%</td></tr><tr><td><strong>Gemini 1.5 Pro</strong></td><td>80.08%</td><td>81.90%</td><td>92.50%</td><td>71.90%</td><td>84.00%</td><td>91.70%</td><td>58.50%</td></tr><tr><td><strong>Gemini Ultra</strong></td><td>79.52%</td><td>83.70%</td><td>87.80%</td><td>74.40%</td><td>83.60%</td><td>94.40%</td><td>53.20%</td></tr><tr><td><strong>GPT-4</strong></td><td>79.45%</td><td>86.40%</td><td>95.30%</td><td>67.00%</td><td>83.10%</td><td>92.00%</td><td>52.90%</td></tr><tr><td><strong>Llama 3 Instruct - 70B</strong></td><td>79.23%</td><td>82.00%</td><td>87.00%</td><td>81.70%</td><td>81.30%</td><td>93.00%</td><td>50.40%</td></tr><tr><td><strong>Claude 3 Sonnet</strong></td><td>76.55%</td><td>79.00%</td><td>89.00%</td><td>73.00%</td><td>82.90%</td><td>92.30%</td><td>43.10%</td></tr><tr><td><strong>Claude 3 Haiku</strong></td><td>73.08%</td><td>75.20%</td><td>85.90%</td><td>75.90%</td><td>73.70%</td><td>88.90%</td><td>38.90%</td></tr><tr><td><strong>Gemini Pro</strong></td><td>68.28%</td><td>71.80%</td><td>84.70%</td><td>67.70%</td><td>75.00%</td><td>77.90%</td><td>32.60%</td></tr><tr><td><strong>GPT-3.5</strong></td><td>65.46%</td><td>70.00%</td><td>85.50%</td><td>48.10%</td><td>66.60%</td><td>57.10%</td><td>34.10%</td></tr><tr><td><strong>Mixtral 8x7B</strong></td><td>59.79%</td><td>70.60%</td><td>84.40%</td><td>40.20%</td><td>60.76%</td><td>74.40%</td><td>28.40%</td></tr><tr><td><strong>Llama 3 Instruct - 8B</strong></td><td>-</td><td>68.40%</td><td>-</td><td>62.00%</td><td>61.00%</td><td>79.60%</td><td>30.00%</td></tr></tbody></table>


# Foundation Models

Training Foundation Models

Foundation models are are trained on broad data, generally using self-supervision at scale, which allows them to be adapted or fine-tuned to a wide range of tasks.&#x20;

{% embed url="<https://arxiv.org/abs/2108.07258>" %}
Foundation Models
{% endembed %}

The paper below gives an excellent overview of the history of large language models:

{% embed url="<https://arxiv.org/abs/2307.06435>" %}
A comprehensive and up to date review of large language models
{% endembed %}

<mark style="color:green;">**The Evolution of AI: From Machine Learning to Foundation Models**</mark>

The journey of AI has been a tale of increasing dependence on data-driven learning.&#x20;

From the traditional machine learning era, where predictive models were trained on historical data, we've now entered the deep learning phase, where models learn and derive high-level features from massive datasets.

This evolution is marked by the rise of deep neural networks and self-supervised learning, relying heavily on large datasets, computational power, and the Transformer model architecture.

<mark style="color:green;">**The Impacts and Risks of Foundation Models**</mark>

While foundation models bring homogenisation in research and new abilities like in-context learning (as seen in GPT-3), they also come with their share of risks. These include the amplification of biases and a dependency on a few models, which necessitated a careful approach in their development and management.

#### <mark style="color:green;">The Name Game: Why "Foundation Models"?</mark>

The term "foundation model" aptly captures the paradigm shift in AI.&#x20;

It signifies a model class that's fundamental, adaptable, and a bedrock for further applications, emphasizing their architectural significance and the sociological impact they've had on AI research and deployment.

### <mark style="color:purple;">The Social Impact and Ecosystem of Foundation Models</mark>

<mark style="color:green;">**Real-World Integration**</mark>

Foundation models have been integrated into real-world systems, like Google Search. This integration has demanded an examination of their social impacts, including fairness, economic and environmental concerns, and ethical implications.

<mark style="color:green;">**The Ecosystem Approach**</mark>

Foundation models are part of a larger AI ecosystem that includes data creation, curation, training, adaptation, and deployment. Each stage has its impact and challenges, highlighting the importance of considering the entire pipeline in understanding and managing these models.

#### <mark style="color:green;">Looking Ahead: The Future of Foundation Models</mark>

As we look to the future, foundation models are poised for further evolution. They are likely to see increased integration across various sectors and will necessitate robust frameworks for their responsible management. The potential for these models to revolutionise technology and society is immense, but so is the need for careful consideration of the societal and ethical challenges they bring.

### <mark style="color:purple;">Conclusion: Embracing the New Era of AI</mark>

Foundation models are a groundbreaking development in AI, offering unmatched versatility and potential.&#x20;

However, their success and beneficial integration into society depend on our collective understanding, responsible management, and thoughtful consideration of their impacts. As we embrace this new era of AI, let's do so with both excitement and caution, ensuring that these powerful tools are used for the greater good.


# LLama 2 - Analysis

Meta introduced Llama 2 during June 2023

The release of LLama 2 a landmark moment, a powerful open-source large language model, available for both research and commercial use at no cost.&#x20;

Llama 2's is the next iteration of LLama 1, a model released January 2023.   Llama 2, available for free, comes with model weights and starting code for both the pre-trained and conversational fine-tuned versions.&#x20;

{% embed url="<https://arxiv.org/abs/2307.09288>" %}
Llaam2 paper
{% endembed %}

Llama 2 is a collection of <mark style="color:yellow;">pretrained and fine-tuned large language models (LLMs)</mark>, ranging from <mark style="color:yellow;">7 billion to 70 billion parameters</mark>. The models, specifically optimised for dialogue use cases, are known as Llama 2-Chat.&#x20;

### <mark style="color:purple;">Pre Training Process</mark>

The pretraining process for the Llama 2 models incorporated several enhancements over its predecessor, Llama 1.&#x20;

<mark style="color:green;">**Pretraining Data and Sources**</mark>

For the Llama 2 models, a new mix of publicly available online data sources was curated, explicitly excluding data from Meta’s products or services and removing content from sites known for containing personal information.&#x20;

The training involved 2 trillion tokens, chosen for their balance of performance and cost-effectiveness.&#x20;

This data selection was also aimed at improving model knowledge and reducing inaccurate predictions or 'hallucinations.'  Comprehensive pretraining data investigations were conducted to understand the potential capabilities and limitations of the models.

### <mark style="color:purple;">**Model Architecture and Training Details**</mark>

The core architecture of Llama 2 models is based on the standard transformer architecture, as defined by Vaswani et al. (2017).&#x20;

Key features of this architecture include:

* <mark style="color:green;">**Pre-Normalization**</mark><mark style="color:green;">:</mark> Utilizing RMSNorm (Zhang and Sennrich, 2019) for stabilising the training process.
* <mark style="color:green;">**SwiGLU Activation Function**</mark><mark style="color:green;">:</mark> Adopted from Shazeer (2020), enhancing the model's capability to capture complex patterns.
* <mark style="color:green;">**Rotary Positional Embeddings (RoPE)**</mark><mark style="color:green;">:</mark> As proposed by Su et al. (2022), which helps the model understand the order of input tokens better.
* <mark style="color:green;">**Grouped-Query Attention (GQA)**</mark><mark style="color:green;">:</mark> This is a significant architectural difference from Llama 1, improving inference scalability for larger models. GQA enables the model to process inputs more efficiently by grouping queries, which is especially beneficial for models with a high number of parameters.

Additionally, Llama 2 models have *<mark style="color:yellow;">**doubled the context length**</mark>* compared to Llama 1, allowing them to consider larger segments of text during training, thereby capturing more extensive contextual information.

### <mark style="color:purple;">**Hyperparameters and Training Optimisation**</mark>

The training used the AdamW optimizer (Loshchilov and Hutter, 2017), with specific settings for beta values, epsilon, learning rate schedule (cosine learning rate with warmup and decay), weight decay, and gradient clipping. These hyperparameters were fine-tuned to optimise the training process and achieve the desired model performance.

<mark style="color:green;">**Tokenizer**</mark><mark style="color:green;">:</mark> The tokenizer employed for Llama 2 is consistent with the one used in Llama 1, using a bytepair encoding (BPE) algorithm (Sennrich et al., 2016) as implemented in SentencePiece (Kudo and Richardson, 2018).  This tokenizer breaks down text into a set of 32,000 tokens, including individual digits and bytes for decomposing unknown UTF-8 characters, aiding in the effective processing of diverse linguistic inputs.

<mark style="color:green;">**Comparative Analysis and Training Efficiency**</mark><mark style="color:green;">:</mark> A comparative analysis of Llama 2 models against Llama 1 models reveals the advancements in token counts, context length, and the introduction of GQA for larger models. The training loss data for Llama 2 models indicates no signs of saturation even after processing 2 trillion tokens, suggesting the potential for further model improvement.

### <mark style="color:purple;">Fine Tuning Phase</mark>

\
The fine-tuning process of Llama 2-Chat involved a combination of <mark style="color:yellow;">supervised fine-tuning (SFT)</mark> and <mark style="color:yellow;">Reinforcement Learning with Human Feedback (RLHF)</mark>, supplemented by a novel technique known as <mark style="color:yellow;">Ghost Attention (GAtt)</mark>.&#x20;

This process, which demands significant computational and annotation resources, is aimed at aligning the model for specific use cases, *<mark style="color:green;">**particularly dialogue interactions.**</mark>*

### <mark style="color:purple;">**Supervised Fine-Tuning (SFT)**</mark>

* <mark style="color:green;">**Initial Data**</mark><mark style="color:green;">:</mark> The process began with publicly available instruction tuning data, which serves as a foundation for further fine-tuning.
* <mark style="color:green;">**Quality of Data**</mark><mark style="color:green;">:</mark> Recognising the limitations of third-party SFT data in terms of diversity and quality, especially for dialogue-style instructions, the team focused on collecting high-quality SFT data. This involved several thousand examples, *<mark style="color:yellow;">**with a total of 27,540 annotations collected.**</mark>*
* <mark style="color:green;">**Fine-Tuning Details**</mark><mark style="color:green;">:</mark> For the actual fine-tuning, the team used a cosine learning rate schedule, an initial learning rate of 2 × 10^−5, weight decay of 0.1, batch size of 64, and a sequence length of 4096 tokens.  The process involved concatenating prompts and responses from the training set, separated by a special token. The training used an autoregressive objective, zeroing out loss on user prompt tokens, which meant only answer tokens were backpropagated. The model underwent fine-tuning for 2 epochs.

### <mark style="color:purple;">**Reinforcement Learning with Human Feedback (RLHF)**</mark>

#### <mark style="color:green;">**Preference-Based Annotation**</mark>

* After Supervised Fine Tuning, the team shifted focus to RLHF, using preference-based annotation. This involved human annotators comparing model-generated samples against human-provided annotations to train a reward model. The output from the SFT model was found to be competitive with the human-written SFT data, indicating the potential for reprioritizing annotation efforts toward RLHF.
* **Human Preference Data Collection**: The RLHF involved collecting human preference data, where annotators chose between two model responses to a prompt, providing feedback on which response was preferable and why. This data was used to train two separate reward models, one optimized for helpfulness and the other for safety.

#### <mark style="color:green;">**Ghost Attention (GAtt)**</mark>

* **Dialogue Flow Control**: The GAtt technique, introduced in the fine-tuning process, was found to be effective in controlling dialogue flow over multiple turns, enhancing the coherence and consistency of the model’s responses in extended dialogues.

### <mark style="color:purple;">Performance</mark>

Comparative evaluations reveal that Llama 2-Chat models demonstrate superior performance against various benchmarks and rival models.&#x20;

They outperform open-source models in single-turn and multi-turn prompts, and show competitive results against closed-source models like ChatGPT.&#x20;

The 34B version of Llama 2-Chat, for instance, exhibited a win rate of over 75% against similar-sized models, and the 70B model surpassed the PaLM-bison chat model in performance.

### <mark style="color:purple;">Safety</mark>

The development of Llama 2 involved a  process to ensure safety and responsibility, particularly during its pretraining phase.

This process was pivotal in understanding the content and implications of the pretraining data, crucial for identifying potential biases and downstream issues that might arise.

<mark style="color:green;">**Steps for Responsible Pretraining**</mark>

Meta's standard privacy and legal review processes were rigorously followed for each dataset used in training. Notably, *<mark style="color:yellow;">**no Meta user data were included in the training process**</mark>*. The team also excluded data from sites with high volumes of personal information to protect individual privacy.

<mark style="color:green;">**Data Toxicity Measurement**</mark>

Toxicity in the pretraining data was assessed using a HateBERT classifier fine-tuned on the ToxiGen dataset. This evaluation showed that a small percentage of the pretraining data contained toxic elements.&#x20;

However, the decision not to overly scrub the data was made to ensure broader applicability of Llama 2, including tasks like hate speech detection.

<mark style="color:green;">**Safety Benchmarks Evaluation**</mark>

Llama 2 was tested against several automatic safety benchmarks to assess its truthfulness, toxicity, and bias.&#x20;

These benchmarks provided insights into the model's ability to produce reliable, non-toxic content and its propensity to reproduce social biases.&#x20;

The evaluations showed that Llama 2 performed variably across different metrics, with some increase in toxicity for larger models, likely due to the larger pretraining data or different dataset mixes.


# Analysis of Llama 3

### <mark style="color:purple;">Key Points</mark>

### <mark style="color:green;">Model Architecture</mark>

* Llama 3 is a transformer-based model that uses a series of transformer blocks with <mark style="color:blue;">**grouped query attention**</mark> instead of regular attention.
* Grouped query attention allows multiple queries to be mapped to a single key-value pair, enabling more efficient processing of information.
* The model architecture consists of an input layer, followed by a sequence of transformer blocks, and an output head.

### <mark style="color:green;">Tokenization</mark>

* Llama 3 employs the <mark style="color:blue;">tiktoken tokenizer</mark>, which is the same tokenizer used by GPT-4.
* The tokenizer has a vocabulary size of 128,000 tokens, enabling efficient encoding of language.
* Tokenization is a crucial step in converting text into integers, which are then transformed into tensors that serve as input to the model.

<details>

<summary><mark style="color:blue;"><strong>Tiktoken: A Fast and Efficient Tokenizer for OpenAI Models</strong></mark></summary>

Tiktoken is a fast and efficient open-source tokenizer developed by OpenAI.&#x20;

It is designed to work seamlessly with OpenAI's language models, providing a reliable and consistent way to tokenize text data.

<mark style="color:green;">**Key Features**</mark>

1. Speed: Tiktoken is known for its exceptional speed, outperforming other comparable open-source tokenizers by a factor of 3 to 6 times.
2. Encoding and Decoding: Tiktoken provides simple methods to encode text into token integers and decode token integers back into text. The encode() method converts a text string into a list of token integers, while the decode() method converts a list of token integers back into a string.
3. Token Counting: Tiktoken makes it easy to count the number of tokens in a text string, which is crucial for determining whether a string is too long for a model to process and estimating the cost of an API call.
4. Multilingual Support: Tiktoken can handle text in various languages, making it suitable for multilingual applications.
5. Byte-Level Decoding: Tiktoken offers a decode\_single\_token\_bytes() method that safely converts a single integer token to the bytes it represents, preventing any loss of information that may occur when decoding individual tokens.

</details>

### <mark style="color:green;">Training Data</mark>

* The Llama 3 model was trained on a massive dataset consisting of 15 trillion tokens, which is 7 times larger than the dataset used for Llama 2.
* The training data was collected from publicly available sources and underwent extensive filtering and quality assurance processes.
* To support multilingual use cases, 5% of the training data consists of non-English text from over 30 languages.

### <mark style="color:green;">Scaling and Parallelization</mark>

* The developers of Llama 3 made significant improvements in scaling and parallelisation techniques to efficiently train the model.
* They combined data parallelisation, model parallelisation, and pipeline parallelisation to distribute the workload across multiple GPUs.
* Custom-built GPU clusters with over 24,000 GPUs were utilized to maximise training efficiency.

### <mark style="color:green;">Reinforcement Learning Techniques</mark>

* Llama 3 incorporates reinforcement learning techniques such as Supervised Fine-Tuning (SFT), Proximal Policy Optimization (PPO), and Direct Policy Optimization (DPO).
* These techniques help improve the model's performance on reasoning and coding tasks by learning to select the correct answers and reasoning traces.

### <mark style="color:green;">Transformer Block</mark>

* The core component of Llama 3 is the transformer block, which consists of an attention mechanism followed by a feed-forward network (FFN).
* The transformer block is applied repeatedly, with the number of iterations determined by the number of layers specified in the model configuration.
* Residual connections and layer normalization are used to stabilize training and improve the flow of information through the network.

### <mark style="color:green;">Attention Mechanism</mark>

* Llama 3 employs grouped query attention, which is considered the best form of attention currently available.
* Grouped query attention allows multiple queries to be mapped to a single key-value pair, enabling more efficient processing of information.
* The attention mechanism computes the similarity between queries and keys, applies a softmax function to obtain attention weights, and then multiplies the weights with the corresponding values.

### <mark style="color:green;">Output Head</mark>

* The output head of Llama 3 consists of a layer normalization, a linear transformation to the vocabulary size, and a softmax function.
* The output head takes the final hidden states from the transformer blocks and projects them into the vocabulary space.
* The softmax function converts the logits into probability distributions over the vocabulary, allowing the model to generate text.

### <mark style="color:green;">Residual Connections and Layer Normalization</mark>

* Residual connections, also known as skip connections, are used to add the input of a transformer block to its output, facilitating the flow of information and gradients through the network.
* Layer normalization is applied to the inputs and outputs of the transformer blocks to normalize the activations and stabilize training.
* The combination of residual connections and layer normalization helps alleviate the vanishing gradient problem and enables the training of deeper networks.

### <mark style="color:green;">Tokenizer Decode</mark>

* The output of the model, which is a probability distribution over the vocabulary, is passed through the tokenizer's decode function to convert the predicted tokens back into human-readable text.
* The tokenizer's decode function maps the integer token IDs to their corresponding subwords or characters, reconstructing the generated text.

### <mark style="color:purple;">Conclusion</mark>

Llama 3 represents a significant advancement in open-source language models, showcasing state-of-the-art performance across various benchmarks.&#x20;

The model's architecture, training data, and optimisation techniques contribute to its impressive capabilities.&#x20;

The use of grouped query attention, reinforcement learning techniques, and efficient parallelization strategies enables Llama 3 to process and generate high-quality text effectively.


# Llama 3.1 series

### Overview

<mark style="color:blue;">**Llama 3.1**</mark> is a collection of multilingual large language models (LLMs) developed by Meta, available in 8B, 70B, and <mark style="color:yellow;">405B</mark> parameter sizes.&#x20;

These models are designed for both text input and output, with a focus on multilingual dialogue use cases. Llama 3.1 stands out for its architectural advancements, extensive training data, and support for a wide array of languages.

### Architecture and Training

Llama 3.1 uses an optimised transformer architecture, employing auto-regressive language modeling.&#x20;

The model incorporates <mark style="color:yellow;">**Grouped-Query Attention (GQA)**</mark> to enhance inference scalability and is trained on over 15 trillion tokens of multilingual data.

It supports a context length of up to 128,000 tokens, allowing for the processing of extensive text inputs.&#x20;

The model was trained using <mark style="color:yellow;">**supervised fine-tuning (SFT)**</mark> and <mark style="color:yellow;">**reinforcement learning with human feedback (RLHF)**</mark>, ensuring high performance across diverse tasks.

### Supported Languages

Llama 3.1 supports a broad range of languages, including:

* English
* German
* French
* Italian
* Portuguese
* Hindi
* Spanish
* Thai

This multilingual capability makes Llama 3.1 versatile for various global applications.

### Performance Comparison

<table data-full-width="true"><thead><tr><th width="175">Benchmark</th><th>Metric</th><th width="123">Llama 3 8B</th><th>Llama 3.1 8B</th><th>Llama 3 70B</th><th>Llama 3.1 70B</th><th>Llama 3.1 405B</th></tr></thead><tbody><tr><td>MMLU (5-shot)</td><td>macro_avg/acc</td><td>68.5</td><td>69.4</td><td>82.0</td><td>83.6</td><td>87.3</td></tr><tr><td>MMLU (CoT, 0-shot)</td><td>macro_avg/acc</td><td>65.3</td><td>73.0</td><td>80.9</td><td>86.0</td><td>88.6</td></tr><tr><td>MMLU-Pro (CoT, 5-shot)</td><td>macro_avg/acc</td><td>45.5</td><td>48.3</td><td>63.4</td><td>66.4</td><td>73.3</td></tr><tr><td>ARC-Challenge (0-shot)</td><td>acc</td><td>82.4</td><td>83.4</td><td>94.4</td><td>94.8</td><td>96.9</td></tr><tr><td>HumanEval (0-shot)</td><td>pass@1</td><td>60.4</td><td>72.6</td><td>81.7</td><td>80.5</td><td>89.0</td></tr><tr><td>GSM-8K (CoT, 8-shot)</td><td>em_maj1@1</td><td>80.6</td><td>84.5</td><td>93.0</td><td>95.1</td><td>96.8</td></tr><tr><td>MATH (CoT, 0-shot)</td><td>final_em</td><td>29.1</td><td>51.9</td><td>51.0</td><td>68.0</td><td>73.8</td></tr><tr><td>API-Bank (0-shot)</td><td>acc</td><td>48.3</td><td>82.6</td><td>85.1</td><td>90.0</td><td>92.0</td></tr><tr><td>Gorilla Benchmark API Bench (0-shot)</td><td>acc</td><td>1.7</td><td>8.2</td><td>14.7</td><td>29.7</td><td>35.3</td></tr><tr><td>Multilingual MGSM (CoT, 0-shot)</td><td>em</td><td>-</td><td>68.9</td><td>-</td><td>86.9</td><td>91.6</td></tr></tbody></table>

### Benchmark Metrics Overview

**Macro Average Accuracy (MMLU, MMLU-Pro)**

* **Definition**: The macro-average accuracy across multiple subjects in the MMLU (Massive Multitask Language Understanding) benchmark.
* **Purpose**: Represents the average performance across various domains, giving equal weight to each subject regardless of the number of questions.

**Accuracy (ARC-Challenge, API-Bank, Gorilla Benchmark API Bench)**

* **Definition**: The accuracy score, representing the proportion of correct answers out of all questions or tasks.
* **Purpose**: Measures the overall correctness of the model's responses in these specific benchmarks.

**Pass\@1 (HumanEval)**

* **Definition**: The percentage of coding problems that the model solved correctly on the first attempt.
* **Purpose**: Used for the HumanEval benchmark, this metric tests the model's code generation capabilities.

**EM\_Maj1\@1 (GSM-8K)**

* **Definition**: EM likely stands for "Exact Match."
* **Purpose**: This metric is used for the GSM-8K benchmark, which tests grade-school level math problem-solving. The "maj1\@1" suggests a majority voting scheme with one attempt.

**Final Exact Match (MATH)**

* **Definition**: The final exact match (EM) accuracy in the MATH benchmark.
* **Purpose**: Tests advanced mathematical problem-solving abilities, focusing on the correctness of the final answer.

**Exact Match (Multilingual MGSM)**

* **Definition**: EM likely stands for "Exact Match."
* **Purpose**: Used for the Multilingual MGSM (Multilingual Grade School Math) benchmark, this metric tests math problem-solving across different languages.

**Note**: For all these metrics, <mark style="color:yellow;">higher percentages indicate better performance</mark>.&#x20;

The "CoT" (Chain of Thought) notation in some benchmarks signifies that the model was prompted to show its reasoning process, not just the final answer.

### Unique Features and Capabilities

* **Long Context Window**: Supports up to 128,000 tokens.
* **Multilingual Input and Output**: Handles multiple languages effectively.
* **Tool Integration**: Capable of integrating with third-party tools.
* **Improved Safety**: Enhanced refusal handling and safety features.

### Comparison to Previous Versions

Llama 3.1 shows consistent improvements over Llama 3 across various benchmarks:

* **MMLU (5-shot)**: <mark style="color:yellow;">83.6%</mark> for Llama 3.1 70B Instruct vs. <mark style="color:yellow;">82.0%</mark> for Llama 3 70B Instruct.
* **MMLU (Chain of Thought, 0-shot)**: <mark style="color:yellow;">86.0%</mark> for Llama 3.1 70B Instruct vs. <mark style="color:yellow;">80.9%</mark> for Llama 3 70B Instruct.
* **HumanEval (0-shot)**: Slight decrease to <mark style="color:yellow;">80.5%</mark> for Llama 3.1 70B Instruct vs. <mark style="color:yellow;">81.7%</mark> for Llama 3 70B Instruct.

### Efficiency Considerations

Llama 3.1 employs Grouped-Query Attention (GQA) to improve inference scalability.&#x20;

The availability of different model sizes (8B, 70B, 405B) allows flexibility in deployment based on resource constraints. The model also supports various fine-tuning techniques like LoRA and QLoRA, enhancing efficiency for specific tasks.

### Fine-Tuning, Quantization, and Prompting

#### Fine-Tuning

Fine-tuning is the process of adapting a pre-trained model to a specific task or dataset.

For Llama 3.1, there are several approaches:

* **Full Parameter Fine-Tuning**: Adjusts all model parameters but is resource-intensive.
* **PEFT (Parameter Efficient Fine Tuning)**:
  * **LoRA (Low Rank Adaptation)**: Uses 8-bit quantized weights.
  * **QLoRA (Quantized LoRA)**: Uses 4-bit quantized weights, requiring even less memory.
* **Tools and Libraries**:
  * `llama-recipes`: Provides scripts for different fine-tuning methods.
  * `torchtune`: Supports the entire fine-tuning lifecycle, including multi-GPU training.
  * `Hugging Face PEFT`: Offers easy-to-use scripts for LoRA fine-tuning.
  * `Axolotl`: An open-source library for streamlined fine-tuning.

### Quantization

Quantization reduces computational and memory requirements by representing weights and activations with lower precision data. For Llama 3.1:

* **PyTorch Quantization Modes**:
  * Post-Training Dynamic Quantization
  * Post-Training Static Quantization
  * Quantization Aware Training (QAT)
* **Tools and Libraries**:
  * `TorchAO`: Offers various quantization methods, including autoquantization.
  * `Hugging Face Transformers`: Supports multiple quantization techniques.
  * `Quanto`: A versatile PyTorch quantization toolkit.
  * `AQLM (Additive Quantization of Language Models)`
  * `AWQ (Activation-aware Weight Quantization)`
  * `AutoGPTQ`: Implements the GPTQ algorithm for post-training quantization.
  * `BitsAndBytes`: Supports 8-bit and 4-bit quantization.

### Prompting

Prompting involves crafting input text to guide the model's output. Key techniques for Llama 3.1 include:

* **Crafting Effective Prompts**:
  * Be clear and concise.
  * Use specific examples.
  * Vary the prompts.
  * Test and refine.
  * Use feedback.
* **Explicit Instructions**: Provide detailed guidelines for better results.
* **Stylization**: Specify the desired style or tone of the response.
* **Formatting**: Request specific output formats (e.g., bullet points, JSON).
* **Restrictions**: Set constraints on the model's responses.
* **Zero-Shot and Few-Shot Learning**: Provide examples to guide the model's understanding.
* **Role-Based Prompts**: Frame the prompt from a specific perspective.
* **Chain of Thought**: Guide the model's reasoning process step-by-step.
* **Self-Consistency**: Generate multiple responses and select the most frequent answer.
* **Retrieval-Augmented Generation (RAG)**: Incorporate external information into prompts.
* **Program-Aided Language Models**: Use code generation for calculations.
* **Techniques to Reduce Hallucinations**: Minimize extraneous tokens.

These techniques allow developers to optimize Llama 3.1's performance, efficiency, and output quality for various applications.

### Key Features

* **Prompt Format**: Llama 3.1 uses a specific prompt format with special tokens to structure interactions.
* **Multiple Roles**: Supports four roles - system, user, assistant, and ipython (for tool interactions).
* **Tool Calling**: The model can integrate with external tools and generate appropriate function calls.
* **Customizable Prompts**: Users can define custom formats for tool interactions.

### How to Use

#### Basic Interaction

* Start with the `<<start>>` token.
* Use role headers like `<<user>>` to denote different parts of the conversation.
* End turns with `<<end>>`.

#### System Instructions

* Set up the context, rules, and available tools in the system prompt.
* Example: `<<system>> Environment: ipython Tools: brave_search, wolfram_alpha You are a helpful assistant. <<end>>`

#### User Queries

* Format user messages with appropriate headers.
* Example: `<<user>> What is the weather in San Francisco? <<end>>`

#### Tool Calling

* Built-in tools (`brave_search`, `wolfram_alpha`, `code_interpreter`) can be activated in the system prompt.
* Custom tools can be defined in JSON format.
* The model generates tool calls in specified formats (Python or JSON).

#### Multi-Turn Conversations

* Continue the conversation by alternating user and assistant roles.
* For tool interactions, use the ipython role to provide tool outputs back to the model.

#### Custom Formats

* You can define custom formats for tool calls in the system prompt.
* Example: Using tags with JSON parameters.

#### Response Handling

* The model uses `<<continue>>` for multi-step reasoning (expecting tool output).
* It uses `<<stop>>` to signal the end of a complete response.

#### Key Points

* The model doesn't execute tool calls; it generates structured output for external execution.
* Developers should test different prompt structures for their specific use cases.

Here's a set of code blocks that demonstrate how to use the Llama 3.1 model based on the documentation provided:

### Basic Interaction Example

```python
# Initialize the model and tokenizer
from transformers import LlamaTokenizer, LlamaForCausalLM

# Load the tokenizer and model
tokenizer = LlamaTokenizer.from_pretrained("meta/llama-3.1-70b")
model = LlamaForCausalLM.from_pretrained("meta/llama-3.1-70b")

# Define a simple user query
input_text = "<<start>> <<user>> What is the capital of France? <<end>>"

# Tokenize the input
inputs = tokenizer(input_text, return_tensors="pt")

# Generate a response
outputs = model.generate(**inputs)

# Decode and print the response
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
```

### System Instructions and Tool Integration

```python
# Example with system instructions and tool integration

input_text = (
    "<<start>> "
    "<<system>> Environment: ipython Tools: wolfram_alpha "
    "You are a helpful assistant. <<end>>"
    "<<user>> What is the current temperature in New York? <<end>>"
)

# Tokenize the input
inputs = tokenizer(input_text, return_tensors="pt")

# Generate a response
outputs = model.generate(**inputs)

# Decode and print the response
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
```

### Custom Tool Calling with JSON Format

```python
# Custom tool integration example

input_text = (
    "<<start>> "
    "<<system>> Tools: {\"weather_tool\": {\"params\": {\"city\": \"New York\"}}} <<end>>"
    "<<user>> Can you tell me the weather in New York? <<end>>"
)

# Tokenize the input
inputs = tokenizer(input_text, return_tensors="pt")

# Generate a response
outputs = model.generate(**inputs)

# Decode and print the response
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
```

### Multi-Turn Conversation Example

```python
# Multi-turn conversation example

input_text = (
    "<<start>> "
    "<<user>> What is the square root of 144? <<end>>"
    "<<assistant>> The square root of 144 is 12. <<end>>"
    "<<user>> Can you tell me the square root of 169? <<end>>"
)

# Tokenize the input
inputs = tokenizer(input_text, return_tensors="pt")

# Generate a response
outputs = model.generate(**inputs)

# Decode and print the response
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
```

### Response Handling with Multi-Step Reasoning

```python
# Example of response handling with multi-step reasoning

input_text = (
    "<<start>> "
    "<<user>> Calculate 25 multiplied by 4, then subtract 10. <<end>>"
    "<<continue>>"  # Expecting tool output for intermediate steps
)

# Tokenize the input
inputs = tokenizer(input_text, return_tensors="pt")

# Generate a response
outputs = model.generate(**inputs)

# Decode and print the response
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
```

These code blocks demonstrate various interactions with the Llama 3.1 model, including basic queries, system instructions, tool integration, multi-turn conversations, and multi-step reasoning. These examples can serve as a foundation for more complex applications using the model.

### Capabilities

### Capabilities of Meta Llama 3 with Retrieval-Augmented Generation (RAG)

### **1. Dynamic Knowledge Integration:**

* **Concept:** RAG enables Meta Llama 3 to dynamically incorporate external information during the inference process. This means that the model is not constrained by its training data, which has a fixed cutoff, but can access and use up-to-date or domain-specific information as needed.
* **Capability:** Meta Llama 3, when augmented with RAG, can answer queries that require current knowledge or insights drawn from specialized datasets. This is particularly valuable for industries that rely on real-time data or have proprietary information that was not included in the model’s training.

<figure><img src="/files/jKobhNIKZ7MgJqrD19s6" alt=""><figcaption></figcaption></figure>

### **2. Contextual Enhancement for Queries:**

* **Concept:** RAG works by retrieving relevant data from external sources and using it to enhance the context of the input query. This additional context helps the model generate more accurate and contextually relevant responses.
* **Capability:** With RAG, Meta Llama 3 can handle complex, context-dependent queries more effectively. By integrating external data into the query, the model can better understand nuances and provide answers that are tailored to the specific context of the inquiry.

### **3. Reduction of Hallucinations:**

* **Concept:** Hallucinations in LLMs refer to instances where the model generates plausible but incorrect or irrelevant information. RAG mitigates this by grounding the model’s responses in real, retrieved data.
* **Capability:** When using RAG, Meta Llama 3 is less likely to produce hallucinated information, especially in areas where its pre-trained knowledge is insufficient. The model’s responses are instead anchored in the specific, retrieved context, leading to more reliable outputs.

### **4. Custom Data Use:**

* **Concept:** Enterprises can leverage RAG to integrate their own proprietary data into the model’s inference process without needing to retrain the model on that data. This allows for the use of sensitive or specialized information while maintaining data security.
* **Capability:** Meta Llama 3, enhanced with RAG, can provide customized responses based on private datasets, making it highly adaptable to specific organizational needs. This capability is particularly beneficial in sectors like finance, healthcare, and legal services, where domain-specific accuracy is crucial.

### **5. Scalability and Flexibility:**

* **Concept:** RAG allows for scalable and flexible deployment by enabling the model to work with large volumes of data and diverse data sources. This can be done without altering the core architecture of the model.
* **Capability:** Meta Llama 3, combined with RAG, can scale to accommodate vast datasets and complex query requirements. This makes it suitable for enterprise-level applications where the model needs to interact with extensive and varied data sources to generate meaningful responses.

### **Implications of RAG for LLM Applications**

RAG transforms the static nature of LLMs by introducing a dynamic, data-driven approach to query handling. In the context of Meta Llama 3, this means:

* **Enhanced Accuracy:** By accessing up-to-date or specialized data, the model can deliver responses that are not only accurate but also relevant to the specific query context.
* **Data Security:** Organizations can safely use their proprietary data with the model, ensuring that sensitive information remains secure while still benefiting from advanced AI capabilities.
* **Versatility:** Meta Llama 3’s ability to work with a wide range of external data sources through RAG makes it adaptable to various industries and use cases, from real-time customer support to domain-specific research assistance.

#### Conclusion

The integration of RAG with Meta Llama 3 significantly extends the model's capabilities, allowing it to deliver more accurate, context-aware, and reliable responses. This enhancement positions Meta Llama 3 as a powerful tool for enterprises looking to leverage the strengths of large language models while addressing their inherent limitations. By enabling the model to dynamically interact with external data sources, RAG transforms Meta Llama 3 into a versatile solution capable of meeting the demands of complex, real-world applications.


# Google Gemini 1.5

#### The Gemini 1.5 family introduces two key models:&#x20;

#### Gemini 1.5 Pro and Gemini 1.5 Flash.&#x20;

#### These models are designed to handle extensive multimodal understanding and leverage an extensive context window, which enables them to process and reason across vast amounts of information.&#x20;

#### The key features of Gemini 1.5 include improved in-context learning, low-resource machine translation, long-document QA, long-context audio recognition, long-context video QA, in-context planning, and unstructured multimodal data analytics.

***

{% embed url="<https://storage.googleapis.com/deepmind-media/gemini/gemini_v1_5_report.pdf>" %}

###

### **Key Features and Performance Metrics**

#### Audio Context and ASR (Automatic Speech Recognition)

**Performance**:

* **Gemini 1.5 Flash** demonstrates increasing efficiency in learning context for Kalamang ASR, with CER (Character Error Rate) improving with the addition of text and audio context.
* Despite being lighter than 1.5 Pro, Gemini 1.5 Flash achieves significant accuracy in ASR tasks.

**Analysis**:

* The model shows a significant improvement in ASR tasks with more context. It performs better in handling text and audio contexts, particularly in complex tasks like speech segmentation and word spelling.

#### Low-Resource Machine Translation

**Performance**:

* **Gemini 1.5 Pro** shows remarkable performance in translating low-resource languages.
* The model's translation accuracy improves consistently with an increasing number of in-context examples, surpassing **GPT-4 Turbo** significantly.

**Analysis**:

* The model's ability to leverage large context windows allows it to perform well in translating languages with limited pre-existing data. This demonstrates its capability to learn and adapt from in-context examples effectively.

#### Long-Document Question Answering (QA)

**Performance**:

* **Gemini 1.5 Pro** outperforms both **Gemini 1.0 Pro** and **GPT-4 Turbo** in answering questions from long documents like "Les Misérables."
* The model maintains context over extensive texts and provides high-quality answers without the need for retrieval-augmented generation (RAG).

**Analysis**:

* The ability to handle entire books as input showcases the model’s robustness in maintaining context over long passages and comprehensively understanding narratives and relationships within texts.

#### Long-Context Audio

**Performance**:

* **Gemini 1.5 Pro** achieves a WER (Word Error Rate) of 5.5% on 15-minute video transcriptions, outperforming models like **USM** and **Whisper**.
* **Gemini 1.5 Flash** performs admirably with a WER of 8.8%, considering its smaller size.

**Analysis**:

* The model’s long-context capabilities allow it to transcribe longer audio segments accurately without additional segmentation and pre-processing.

#### Long-Context Video QA

**Performance**:

* **Gemini 1.5 Pro** achieves state-of-the-art accuracy in long-video QA tasks, significantly outperforming **GPT-4V**.
* The model’s performance improves as more frames are provided, demonstrating its effectiveness in handling extended video contexts.

**Analysis**:

* The model's ability to handle long video contexts and answer questions based on extensive video content makes it highly effective for applications in multimedia analysis and video understanding.

#### In-Context Planning

**Performance**:

* **Gemini 1.5 Pro** outperforms other models in planning tasks expressed in PDDL (Planning Domain Definition Language) and natural language.
* The model's performance improves with more examples, highlighting its effectiveness in in-context learning for planning tasks.

**Analysis**:

* The model's capability to generate plans based on in-context examples showcases its potential in applications requiring strategic reasoning and decision-making.

#### Unstructured Multimodal Data Analytics

**Performance**:

* **Gemini 1.5 Pro** demonstrates superior performance in extracting structured information from unstructured data like images.
* The model's accuracy improves with larger context windows, outperforming **GPT-4 Turbo** and **Claude 3 Opus**.

**Analysis**:

* The model’s ability to process and analyze unstructured data efficiently makes it valuable for tasks involving large datasets of images, conversations, and other non-text data.

***

#### Cost and Usage

**Cost**:

* **Gemini 1.5 Flash** is designed to be more cost-efficient compared to **Gemini 1.5 Pro**, making it suitable for high-volume, high-frequency tasks.

**Usage**:

* The models can be accessed via API, with costs typically based on the number of tokens processed or the duration of usage. Specific pricing details should be obtained from the official pricing page or by contacting the sales team.

***

### Prompt Techniques and Parameter Settings

**Prompt Techniques**:

**Clear and Specific Instructions**:

```python
prompt = "Summarize the following document with a focus on key financial metrics:\n\n[Document Text]"
```

**Contextual Prompts**:

```python
prompt = "Given the previous quarterly reports, analyze the performance trends and highlight significant changes:\n\n[Previous Reports]\n\n[Current Report]"
```

**Multimodal Inputs**:

```python
prompt = "Provide a caption for the following image:\n\n[Image URL or Description]"
```

**Interactive Prompts**:

```python
prompt = "User: What is the weather forecast for tomorrow?\nAssistant: The weather forecast for tomorrow is sunny with a high of 75°F. Do you need information on any other days?\nUser: Yes, what about the day after tomorrow?"
```

### **Parameter Settings**:

**Temperature**: Controls the randomness of the output.

```python
temperature = 0.7
```

**Max Tokens**: Sets the maximum number of tokens to generate.

```python
max_tokens = 150
```

**Top\_p**: Enables nucleus sampling, controlling the diversity of the output.

```python
top_p = 0.9
```

**Frequency\_penalty**: Discourages repetition of the same phrases.

```python
frequency_penalty = 0.5
```

**Presence\_penalty**: Encourages introducing new topics.

```python
presence_penalty = 0.3
```

#### Conclusion

Gemini 1.5 Flash and Pro models represent significant advancements in handling multimodal data, leveraging long-context windows, and performing a variety of high-frequency, high-volume tasks efficiently. Their ability to excel in complex tasks such as machine translation, long-document QA, and video understanding makes them invaluable tools for a wide range of applications in natural language processing and artificial intelligence.


# Platypus: Quick, Cheap, and Powerful Refinement of LLMs

This <mark style="color:blue;">**March 2024**</mark> demonstrates the potential performance of base Large Language Models (LLMs) through parameter-efficient fine-tuning (PEFT) on a curated dataset named Open-Platypus, focusing on the domain of LLMs in customer support settings.&#x20;

The authors provide a context for their work against the backdrop of significant advancements in LLMs, noting the development of models like PaLM, GPT-3, and LLaMa, which emphasise computational efficiency and the movement towards open-source models like BLOOM and Falcon.

{% embed url="<https://arxiv.org/abs/2308.07317>" %}
Platypus: Quick, Cheap, and Powerful Refinement of LLMs
{% endembed %}

The paper discusses various strategies to improve LLM performance, including knowledge distillation, instruction tuning, and the Mixture of Experts approach.&#x20;

These methods aim to enhance the models' efficiency and adaptability across various domains. The authors specifically employ the LoRA methodology, noting its effectiveness in their workflow and its potential for future cost and time reductions in training.

Key contributions of the paper include:

<mark style="color:green;">**Open-Platypus Dataset**</mark>

Open-Platypus is a curated dataset that the team created by selecting a <mark style="color:yellow;">subset from other open datasets.</mark> &#x20;

It integrates <mark style="color:yellow;">11 open-source datasets,</mark> predominantly consisting of human-designed questions, enabling robust performance with minimal fine-tuning time and cost.

The high-quality nature of Open-Platypus has allowed for strong performance and efficiency, demonstrating the importance of targeted and specific datasets in training sophisticated models. The dataset is also released to the public, fostering collaborative improvement.

<mark style="color:green;">**Dataset Optimisation**</mark>

The authors describe their process of similarity exclusion to streamline the dataset by reducing redundancy and a training data filtering process to avoid contamination, ensuring the dataset's integrity and relevance.

<mark style="color:green;">**Fine-tuning and Merging Process**</mark>

They detail their approach to selecting and merging specialised fine-tuned LoRA modules, highlighting the effectiveness of this method in imparting specific domain knowledge while maintaining the benefits of instruction tuning.

This work aims to advance the field by providing an efficient way to enhance LLMs for specific tasks, particularly in customer support, emphasizing the potential of domain-specific datasets and merging techniques to improve model performance while reducing training time and costs.

### <mark style="color:purple;">Dataset Creation</mark>

The paper outlines a detailed process for curating the Open-Platypus dataset, aimed at enhancing the performance of Large Language Models (LLMs), particularly focusing on the STEM domain.&#x20;

Here's a breakdown of the data curation process, demystifying any jargon and explaining technical concepts:

<mark style="color:green;">**Data Selection Criteria**</mark>

The curation process was influenced by several theoretical frameworks and empirical findings suggesting that with minimal yet targeted training data, significant alignment of model outputs can be achieved. The dataset aimed to provide depth in specific areas, ensuring diversity in input prompts while maintaining a manageable size.

<mark style="color:green;">**Open-Platypus Dataset Composition**</mark>

This dataset is an aggregation of 11 open-source datasets, predominantly comprising human-generated questions, with about 10% contributed by an LLM. The focus is on STEM and logic, selecting datasets that offer questions in these domains or filtering broader datasets for relevant content.

<mark style="color:green;">**Instruction Tuning**</mark>

To enhance the dataset's effectiveness, an instruction-tuning format was employed where each data point includes an instruction, input, and output. This format is particularly useful for creating structured and consistent training material for the LLM.

<mark style="color:green;">**De-duplication and Similarity Removal**</mark>

To prevent the model from simply memorizing answers, a de-duplication process was implemented. This involved removing exact duplicates and questions with a high degree of similarity (measured by cosine similarity) to others in the dataset. This step ensures that the training data encourages the model to learn underlying patterns and logic rather than memorizing specific answers.

<mark style="color:green;">**Contamination Check**</mark>

A critical part of the curation process involved ensuring that the training data did not contain any questions from benchmark test sets. This prevents the model from giving the illusion of high performance by simply recalling answers to known questions.

<mark style="color:green;">**Fine-tuning and Merging Process**</mark>

The paper also details the use of Low Rank Approximation (LoRA) for fine-tuning the models, which is a technique that adjusts a small set of parameters in the model, making the training process more efficient and cost-effective. The fine-tuning process was carefully managed to ensure that the models improved in the target domains without requiring extensive computational resources.

The meticulous curation process of Open-Platypus aims to ensure that the fine-tuned LLMs are not only effective in their domain-specific tasks but also efficient in terms of training requirements, thereby addressing a critical aspect of AI model development.

### <mark style="color:purple;">Results</mark>

\ <mark style="color:green;">**Performance Overview**</mark>

The Platypus2-70B instruct variant achieved the top position on the Hugging Face Open LLM Leader board with an impressive average score of 73.13, showcasing its superior performance among other models. The Stable-Platypus2-13B model was highlighted as the leading 13 billion parameter model with an average score of 63.96.

<mark style="color:green;">**Model Merging and Fine-tuning**</mark>

The study explored the effects of merging different models (broad and niche) and the benefits of fine-tuning using the Open-Platypus dataset.  The results showed that the fine-tuned models outperformed the base models, particularly in the ARC and TruthfulQA benchmarks, demonstrating the effectiveness of the merging and fine-tuning strategy.

<mark style="color:green;">**Impact on Various Benchmarks**</mark>

The fine-tuned models showed varied performance across different benchmark tests. For example, the Camel-Platypus2-70B model significantly improved in the ARC-Challenge, whereas the Dolphin-Platypus2-70B merge did not surpass the performance of the base and adapter models. This indicates that the merging process's success can vary based on the models and datasets involved.

<mark style="color:green;">**Domain-Specific Performance**</mark>

The effectiveness of the fine-tuned models was domain-specific. For instance, in the machine learning domain, the Camel-Platypus2-70B model showed a remarkable improvement, suggesting that the choice of model for merging is crucial depending on the domain or task at hand.

<mark style="color:green;">**Notable Improvements and Declines**</mark>

The analysis highlighted significant improvements and declines in different domains. For example, the Camel-Platypus2-70B model excelled in the ARC-Challenge, while several models showed notable declines in the college physics test, indicating potential compatibility issues or limitations in certain domains.

<mark style="color:green;">**Insights into Merging Strategy**</mark>

The results provided insights into the merging strategy's complexity, showing that not all merges lead to superior models. The variability in performance across different benchmarks suggests that careful consideration is required when selecting models for merging, especially when targeting specific domains or tasks.

Overall, the results emphasize the potential of fine-tuning and merging strategies to enhance LLMs' performance, demonstrating significant improvements in specific domains while also highlighting the importance of domain-specific evaluations and the complexities involved in the model merging process.

### <mark style="color:purple;">Conclusion</mark>

This paper discusses the enhancement of Large Language Models (LLMs) through fine-tuning using the Open-Platypus dataset and explores the potential benefits of merging smaller, efficient models with the precision of individual adapters.&#x20;

It highlights the success of these fine-tuned models in specific tasks and suggests that future work could explore integrating various datasets and methodologies like QLoRA to improve model performance.&#x20;

The paper acknowledges the limitations of the Platypus model, such as its static knowledge base, potential bias, and its primary focus on English-language data.&#x20;

It stresses the importance of responsible use and the need for further safety testing before deploying the model in applications. The paper also notes the significance of ensuring no contamination between training and benchmark test sets to maintain the integrity of the model's performance. Lastly, it acknowledges the contributions of Hugging Face and Meta AI in supporting the development and evaluation of LLMs.


# Mixtral of Experts

The paper introduces Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model, showcasing advancements in the field of AI language processing.&#x20;

{% embed url="<https://arxiv.org/abs/2401.04088>" %}
Mixtral of Experts
{% endembed %}

### <mark style="color:purple;">Mixtral 8x7B Overview</mark>

<mark style="color:green;">**Architecture**</mark><mark style="color:green;">:</mark> Mixtral 8x7B is built on the same architecture as Mistral 7B but includes a unique feature where each of its layers contains 8 feedforward blocks, known as experts.

<mark style="color:green;">**Expert Selection**</mark>: For every token in the input, a router network at each layer selects two experts to process the token and combine their outputs. This selective process allows the model to dynamically choose which parts of its network to use for each token.

<mark style="color:green;">**Parameter Efficiency**</mark><mark style="color:green;">:</mark> While the total parameter count for the model is 47 billion, any given token is processed using only 13 billion active parameters, enhancing the model's efficiency during inference.

<mark style="color:green;">**Training and Performance**</mark><mark style="color:green;">:</mark> Mixtral was trained with a large context size of 32,000 tokens and shows superior or comparable performance to Llama 2 70B and GPT-3.5 on various benchmarks, particularly excelling in mathematics, code generation, and multilingual tasks.

### <mark style="color:purple;">Analysis of how Mixture of Experts works</mark>

This excellent article from the famous Cameron Wolfe.  Please visit his website here:

{% embed url="<https://cameronrwolfe.me/>" %}
Cameron Wolfe is a generous AI researcher - bridging the gap between academia and practitioners
{% endembed %}

<mark style="color:green;">Here his full article on how Mixture of Experts works:</mark>

{% embed url="<https://substack.com/@cwolferesearch/note/c-51916763>" %}
Mixture of Experts
{% endembed %}

I have paraphrased his article below:

The paper delves into the <mark style="color:blue;">Sparse Mixture of Experts (MoE) layers</mark> in the context of decoder-only transformer architecture, commonly employed in autoregressive large language models (LLMs).&#x20;

It explains how *<mark style="color:yellow;">**MoEs enhance model capacity while maintaining computational efficiency**</mark>* by selectively activating a subset of parameters during the forward pass.

### <mark style="color:purple;">Key Aspects of MoE in LLMs</mark>

<mark style="color:green;">**Expert Architecture:**</mark> Each expert within an MoE layer is a feed-forward neural network with its unique set of parameters, mirroring the architecture of the standard feed-forward sub-layer in traditional transformer models.

<mark style="color:green;">**Routing Mechanism:**</mark> A router processes each token, producing a probability distribution that *<mark style="color:yellow;">**dictates which expert(s) will process the token.**</mark>*  This selective routing significantly increases model capacity without proportionally increasing computational demands.

<mark style="color:green;">**Sparse Activation:**</mark> In each MoE layer, *<mark style="color:yellow;">**only a few experts are activated for each token,**</mark>* reducing the computational cost compared to a model where all parameters are always active.

<mark style="color:green;">**Implementation in Decoder-only Architecture:**</mark> MoEs are integrated into the decoder-only transformer architecture, replacing the standard feed-forward sub-layers with MoE layers. This architecture is prevalent in autoregressive LLMs.

<mark style="color:green;">**Creating an MoE Layer:**</mark> An MoE layer consists of multiple experts, and the layer replaces the conventional feed-forward sub-layer in the transformer block. This setup allows for the model to scale up the number of experts without incurring the full computational cost typically associated with larger models.

<mark style="color:green;">**Routing Details:**</mark> Routing in MoEs usually employs a softmax gating function that converts the outputs of a linear layer into a probability distribution over experts, guiding the token through the most relevant experts based on the routing mechanism's decision.

<mark style="color:green;">**Popularity of MoEs in LLMs:**</mark> MoEs offer a path to scale model capacity - a crucial factor for improving LLM performance. They allow for significant parameter expansion without commensurate increases in computational requirements during training and inference, making large-scale LLMs more feasible and efficient.

<mark style="color:green;">**Example of MoE Implementation:**</mark> An instance is provided where a model named Grok incorporates eight experts per MoE layer, demonstrating how MoEs function in a practical LLM context. Despite having 314B parameters, Grok only activates 25% of these for any given token, illustrating the efficiency gains from sparse activation.

### <mark style="color:purple;">Definitions</mark>

<mark style="color:blue;">**Dense Feed-forward Layers**</mark><mark style="color:blue;">:</mark> These layers are standard components in neural networks where each input is connected to each output by a learned weight. In the context of transformer architectures, dense feed-forward layers are present within each transformer block and process the input sequentially, applying the same set of weights across the entire sequence.

<mark style="color:blue;">**Router**</mark><mark style="color:blue;">:</mark> In the context of a Mixture of Experts (MoE) layer, a router is a mechanism that determines how the input is distributed among the various experts. It takes each token or piece of input data, computes a probability distribution over the experts, and routes the input to the selected experts based on this distribution. The routing process is crucial for managing the computational load and ensuring that the most relevant experts handle each piece of input.

<mark style="color:blue;">**Decoder-only Architecture**</mark><mark style="color:blue;">:</mark> This architecture refers to transformer models that only use the decoder component, omitting the encoder. In autoregressive large language models (LLMs), the decoder-only architecture is prevalent, where each transformer block within the decoder contains layers for masked self-attention and feed-forward processing. The architecture is designed to generate output one token at a time, using previously generated tokens as context.

<mark style="color:blue;">**Sparse Activation**</mark><mark style="color:blue;">:</mark> This concept refers to the activation of only a small subset of the available experts in an MoE layer for any given input. Despite the MoE layer having a large number of experts, sparse activation ensures that only the most relevant experts (determined by the router) are utilized during the forward pass, thus optimizing computational efficiency.

<mark style="color:blue;">**Routing Mechanism**</mark><mark style="color:blue;">:</mark> The routing mechanism in MoE layers, typically a softmax gating function, is responsible for allocating input tokens to different experts. It involves passing the input through a linear layer to produce logits, applying a softmax function to generate a probability distribution, and then using this distribution to select and weigh the contribution of each expert in processing the input.

### <mark style="color:purple;">Summary of Wolfe Analysis</mark>

The incorporation of Sparse Mixture of Experts layers within LLMs enables a substantial increase in model capacity while managing computational expenses.

This methodology allows for the development of more powerful and nuanced language models, potentially unlocking new capabilities in natural language processing and understanding.

### <mark style="color:purple;">Back to Mixtral 8x7B–Instruct</mark>

<mark style="color:green;">**Fine-tuned Model**</mark><mark style="color:green;">:</mark> A variant of Mixtral, known as Mixtral 8x7B–Instruct, is fine-tuned to follow instructions better, demonstrating enhanced performance in human evaluation benchmarks compared to other leading models.

<mark style="color:green;">**Reduced Biases and Balanced Sentiment**</mark><mark style="color:green;">:</mark> The instruction-following version of Mixtral also shows improvements in reducing biases and achieving a more balanced sentiment in its outputs.

### <mark style="color:purple;">Mixtral's Architecture</mark>

<mark style="color:green;">**Transformer Basis**</mark><mark style="color:green;">:</mark> At its core, Mixtral is based on the transformer architecture, known for its efficiency and effectiveness in handling sequence data. Transformers consist of layers with two main sub-blocks: a self-attention mechanism and a feedforward neural network.

<mark style="color:green;">**Modifications**</mark><mark style="color:green;">:</mark> Unlike standard transformers, Mixtral replaces the feedforward blocks with Mixture-of-Expert (MoE) layers. This allows the model to dynamically select which parts of the network to use for processing each token, enhancing its adaptability and efficiency.

<mark style="color:green;">**Context Length**</mark><mark style="color:green;">:</mark> Mixtral is designed to handle a fully dense context length of up to 32,000 tokens, providing it with a substantial lookback capability, beneficial for understanding and generating long sequences of text.

### <mark style="color:purple;">Sparse Mixture of Experts (SMoE)</mark>

<mark style="color:green;">**Expert Networks**</mark><mark style="color:green;">:</mark> In an MoE layer, there are 'n' expert networks ($$E0​,E1​,...,En−1​$$). Each <mark style="color:yellow;">expert is a specialised feedforward network</mark> capable of handling specific types of information or patterns within the data.

<mark style="color:green;">**Gating Network**</mark><mark style="color:green;">:</mark> The gating network <mark style="color:yellow;">decides which experts to engage for processing each token</mark>. It outputs a gating vector $$G(x)$$ which is an n-dimensional vector indicating the relevance of each expert for the current input token $$x$$

<mark style="color:green;">**Output Computation**</mark><mark style="color:green;">:</mark> The output of the MoE layer for an input $$x$$ is a weighted sum of the outputs from the expert networks. Mathematically, it's represented as:

$$
y = \sum\_{i=0}^{n-1} G(x)\_i \cdot E\_i(x)
$$

where $$G(x)i​$$ is the i-th element of the gating vector, and $$Ei​(x)$$ is the output from the i-th expert network.

<mark style="color:green;">**Top-K Gating**</mark><mark style="color:green;">:</mark> The gating mechanism selects the top $$K$$ experts based on the gating vector's values. This selection is made using a <mark style="color:blue;">softmax function</mark> applied to the top-K logits of a linear layer. The mathematical expression for this gating function is:

$$
G(x) = \text{Softmax}(\text{TopK}(x \cdot W\_g))
$$

where $$Wg​$$ is the weight matrix for the gating network.

<mark style="color:green;">**Computational Efficiency**</mark>

By using only the top $$K$$ experts (with  being much smaller than $$n$$), the model reduces the computational load, making it efficient while still leveraging a large parameter space.

#### <mark style="color:green;">Execution and Parallelism</mark>

<mark style="color:blue;">**Efficiency on GPUs**</mark><mark style="color:blue;">:</mark> The MoE layers can be efficiently executed on GPUs, using specialized kernels like Megablocks for sparse matrix operations, enhancing execution speed.

<mark style="color:blue;">**Expert Parallelism**</mark><mark style="color:blue;">:</mark> To scale and distribute the workload, Mixtral employs <mark style="color:blue;">Expert Parallelism</mark>, where *<mark style="color:yellow;">**each expert is processed on a different GPU**</mark>*, allowing parallel processing and reducing computation time.

<mark style="color:blue;">**Load Balancing**</mark><mark style="color:blue;">:</mark> Ensuring an even distribution of workload across GPUs is crucial to prevent bottlenecks and maximize resource utilisation.

In summary, Mixtral's architecture, with its integration of the SMoE approach, represents a significant innovation in transformer-based models, offering a dynamic, efficient, and scalable solution for processing large-scale language data.

### <mark style="color:purple;">Conclusion</mark>

In this study, we unveiled Mixtral 8x7B, a pioneering mixture-of-experts network that achieves state-of-the-art performance among open-source models.&#x20;

Notably, the Mixtral 8x7B Instruct variant surpasses other leading models like Claude-2.1, Gemini Pro, and GPT-3.5 Turbo in human evaluation benchmarks.&#x20;

Remarkably, *<mark style="color:yellow;">**Mixtral achieves this superior performance while using only 13 billion active parameters per token, in stark contrast to its predecessor, Llama 2 70B, which uses 70 billion.**</mark>*&#x20;


# Mixture-of-Agents (MoA)

The Mixture-of-Agents (MoA) methodology leverages the strengths of multiple large language models (LLMs) to enhance performance in natural language understanding and generation tasks.&#x20;

MoA constructs a layered architecture where each layer contains several LLM agents.&#x20;

Each agent uses outputs from the previous layer to generate responses, significantly improving over state-of-the-art models like GPT-4 Omni. MoA achieves a score of 65.1% on AlpacaEval 2.0, outperforming GPT-4 Omni's 57.5%, using only open-source models.

{% embed url="<https://arxiv.org/abs/2406.04692>" %}

LLMs have made significant advancements but face constraints like model size and training data, which are costly to scale.&#x20;

Different LLMs excel in various tasks, raising the question of how to harness their collective expertise. MoA addresses this by leveraging the collaborative strengths of multiple LLMs, where each model improves its responses based on outputs from others, even if the initial outputs are of lower quality.

### <mark style="color:purple;">Methodology</mark>

**Collaborativeness of LLMs**

LLMs generate better responses when referencing outputs from other models. MoA capitalizes on this by using multiple models in a layered architecture:

1. **Layer 1**: Agents generate responses independently.
2. **Layer 2**: Agents refine these responses using outputs from Layer 1.
3. **Layer 3**: Further refinement continues through additional layers.

<figure><img src="/files/fvUj9IYcrRMHWfpVMAy9" alt=""><figcaption><p>Illustration of the Mixture-of-Agents Structure. This example showcases 4 MoA layers with 3 agents in each layer. The agents here can share the same model.</p></figcaption></figure>

**Structure of MoA**

MoA comprises `l` layers, each with `n` LLM agents. The final layer synthesizes these responses into a single high-quality output using an Aggregate-and-Synthesize prompt.

<mark style="color:blue;">**Analogy to Mixture-of-Experts (MoE)**</mark>

MoA extends the MoE concept by operating at the model level, leveraging full-fledged LLMs rather than sub-networks within a single model. This approach eliminates the need for fine-tuning and allows flexibility and scalability with off-the-shelf models.

### <mark style="color:purple;">Evaluation</mark>

**Benchmarks**

MoA was evaluated on AlpacaEval 2.0, MT-Bench, and FLASK benchmarks, demonstrating significant improvements:

* **AlpacaEval 2.0**: Achieved a score of 65.1%, outperforming GPT-4 Omni's 57.5%.
* **MT-Bench**: Secured top positions, even with marginal improvements over already high-performing models.
* **FLASK**: Showed substantial improvements in robustness, correctness, efficiency, factuality, commonsense, and insightfulness.

**Budget and Token Analysis**

MoA is cost-effective, outperforming models like GPT-4 Turbo by approximately 4% while being twice as cost-effective. It also efficiently utilizes computational resources to maximize performance.

#### Key Insights

* **Collaborativeness**: LLMs tend to generate higher quality responses when leveraging outputs from other models.
* **Model Diversity**: Using diverse models in each MoA layer improves performance.
* **Proposer and Aggregator Roles**: Certain models excel in generating reference responses (proposers), while others synthesize these into high-quality outputs (aggregators).

#### Conclusion

The Mixture-of-Agents approach significantly enhances the capabilities of LLMs by leveraging the collective strengths of multiple models.&#x20;

This methodology leads to superior performance, demonstrating the benefits of integrating diverse perspectives from various models. The systematic optimisation of MoA offers a promising direction for future research and development in natural language processing.


# Phi 1.5

The Diminutive Giant: How the Phi Model is Revolutionising AI Accessibility

In a world where bigger has long been synonymous with better, the training of artificial intelligence is witnessing a paradigm shift, a phenomenon I like to call the 'Diminutive Revolution'.&#x20;

{% embed url="<https://arxiv.org/abs/2309.05463>" %}

The Phi model stands at the forefront of this revolution. With just 1.3 billion parameters, it is a David among Goliaths like GPT-3 and GPT-4. Yet, like its biblical counterpart, Phi's size belies its power. This model, small enough to nestle in the palm of your hand via a smartphone, it's a symbol of a new era in AI: one that values efficiency and accessibility over scale.

### <mark style="color:purple;">Emphasising Quality Over Quantity</mark>

The Phi model's performance, despite its smaller size, is a testament to the evolving philosophy in AI development: the supremacy of data quality over quantity.&#x20;

By using a *<mark style="color:yellow;">**meticulously curated, high-quality dataset**</mark>*, Phi achieves feats that rival larger predecessors. This shift towards prioritising data quality over volume is a stride in AI's evolution, echoing a broader understanding that the effectiveness of models is impacted by the quality of their training data.

### <mark style="color:purple;">Innovation in Training: The Synthetic Leap</mark>

A key ingredient in Phi's success is its innovative training methodology, particularly its *<mark style="color:yellow;">**use of synthetic data**</mark>*.&#x20;

This approach, which involves the creation and use of artificial datasets, has enabled the Phi model to specialise and excel in Python coding tasks with remarkable efficiency and accuracy. Such advancements in training methods are paving the way for developing highly functional models that are not only smaller but also specifically tailored for precise applications.

### <mark style="color:purple;">Reshaping the Future of Model Scaling</mark>

Phi's development signals a potential departure from the race to create ever-larger models in the field of AI. Future advancements may increasingly focus on enhancing data quality, refining training techniques, and optimizing model architecture. This approach promises similar, if not superior, outcomes with more manageable and cost-effective models.

### <mark style="color:purple;">The Rise of Specialisation</mark>

Phi embodies a trend towards the *<mark style="color:yellow;">**creation of specialised models designed for specific tasks**</mark>*.&#x20;

Its proficiency in Python coding exemplifies how targeted fine-tuning can yield significant improvements in areas beyond those explicitly featured in its training.&#x20;

This move towards specialization indicates a broader applicability of these techniques across various domains, suggesting an exciting future where AI can be custom-tailored to myriad specific needs.

### <mark style="color:purple;">The Interplay of Model and Data Quality</mark>

The choice of using GPT-4 over GPT-3.5 to generate synthetic data for training Phi underscores a crucial lesson: the quality of both the AI model and the data it consumes is paramount.&#x20;

Lower error rates in GPT-4 generated data have led to significant gains in Phi's performance, highlighting the intricate dance between model and data quality in achieving exceptional AI outcomes.

### <mark style="color:purple;">Wizard Coder: A Case Study in Efficiency</mark>

Another illustration of this trend is the 'Wizard Coder'.&#x20;

Despite having fewer parameters (16 billion) compared to larger models, its success is attributed to training on more complex and challenging examples. This highlights a growing understanding that increasing the depth and complexity of training data can lead to substantial improvements, even in relatively smaller models.

### <mark style="color:purple;">The Cambrian Explosion of Specialised AIs</mark>

We are possibly on the cusp of a 'Cambrian explosion' of specialised AIs, as evidenced by Phi 1.5's specialisation in Python coding.&#x20;

This shift towards creating AI tailored for specific tasks values the quality of task-specific datasets and marks a divergence from the trend of scaling up models for the sake of size.

### <mark style="color:purple;">Navigating the Ethical Maze</mark>

As we marvel at these technological leaps, we must also navigate the ethical labyrinth they present. AI safety, especially concerning biological misuse and the creation of harmful pathogens, is a pressing concern. This calls for focused public messaging and policy considerations to ensure AI's responsible development.

### <mark style="color:purple;">The Accelerated Path to AGI</mark>

Conversations with experts suggest that significant AI advancements, and possibly even the attainment of Artificial General Intelligence (AGI), may be closer than we think.&#x20;

The trajectory of AI development, driven by rapid resource allocation, improvements in data quality, algorithmic advancements, and hardware innovations, points to a future arriving much sooner than anticipated.

In conclusion, the Phi model and its contemporaries are not just technological advancements; they are harbingers of a new AI era.&#x20;

An era where efficiency, specialisation, and ethical consideration take center stage, reshaping our approach to artificial intelligence and its role in our lives. As we stand on this precipice of change, one thing is clear: the future of AI is not just about scaling up; it's about thinking smarter.


# Refining the Art of AI Training: A Deep Dive into Phi 1.5's Innovative Approach

A long list of lessons, tips and tricks from the team that bought us Phi

<mark style="color:green;">**Bridging Efficiency and Capability in Language Models**</mark>

Phi 1.5, a specialised language model for coding, has just <mark style="color:yellow;">1.3 billion parameters</mark>, a notable deviation from the trend of increasingly larger models. This compact size suggests enhanced processing efficiency and reduced resource demands, making it a trailblazer in efficient AI design.

<mark style="color:green;">**Synthetic Data as a Training Game-Changer**</mark>

Phi 1.5 leverages a blend of high-quality textbook data and synthetic data, showcasing a pioneering approach in AI training. This strategy reduces reliance on vast real-world datasets and addresses biases inherent in such data.

<mark style="color:green;">**Revolutionising Training Efficiency**</mark>

Phi's training, completed in just four days using 8 A100 GPUs, marks a leap in training efficiency. This has profound implications for the accessibility and environmental footprint of large language models.

<mark style="color:green;">**Small Size, Big Performance**</mark>

Phi 1.5's small size does not hinder its performance; it achieves a 50% pass-at-one accuracy in human evaluations, challenging the belief that bigger is always better in AI.

<mark style="color:green;">**Prioritising Data Quality**</mark>

The focus on high-quality data during Phi's training highlights the pivotal role of data excellence over sheer volume, potentially reshaping deep learning scaling laws.

<mark style="color:green;">**Curriculum-Based Learning for AI**</mark>

Phi's training employs a curriculum that gradually increases in complexity, mirroring human learning methods. This structured progression could lead to more robust and adaptable AI models.

<mark style="color:green;">**Specialisation in AI: The Code Generation Focus**</mark>

Phi's specialisation in generating Python code from docstrings illustrates the trend towards task-specific language models, moving away from a generalist AI approach.

<mark style="color:green;">**Unpredictability and Versatility in LLMs**</mark>

Phi 1.5's emergent abilities, showing proficiency in tasks beyond its training, highlight the unpredictable nature and versatility of language models.

<mark style="color:green;">**Importance of Contextual Training Data**</mark>

The significance of using self-contained and contextually complete data for training is particularly relevant for models trained on code.

<mark style="color:green;">**Mixture of Experts in AI Models**</mark>

Discussing the mixture of experts approach, as seen in models like GPT-4, provides insights into strategies for enhancing overall AI performance.

<mark style="color:green;">**AI-Powered Data Curation**</mark>

Using a transformer-based classifier to filter code datasets represents an innovative approach to ensuring high-quality training data.

<mark style="color:green;">**Efficient Data Annotation Using AI**</mark>

Employing GPT-4 for dataset annotation presents an efficient alternative to labor-intensive human annotation, addressing ethical concerns.

<mark style="color:green;">**Leveraging Traditional Techniques**</mark>

The use of a random forest classifier for quality assessment exemplifies the effectiveness of combining traditional machine learning methods with modern AI.

<mark style="color:green;">**Encouraging Creativity in AI Training**</mark>

Inducing language models to generate more creative and diverse outputs remains a challenge, especially when using synthetic data.

<mark style="color:green;">**Enhancing Logical Reasoning Through Code Training**</mark>

Training on code not only improves coding logic but also enhances the model’s general logical reasoning skills.

<mark style="color:green;">**Diverse Training Data Generation Techniques**</mark>

The generation process for diverse training data, including topic constraints and target audience variations, aims for content diversity and complexity.

<mark style="color:green;">**Decoder-Only Transformer Architecture**</mark>

Phi's use of a decoder-only transformer, suitable for language generation tasks, contrasts with the encoder-decoder structure common in translation tasks.

<mark style="color:green;">**Flash Attention for Enhanced Memory Efficiency**</mark>

Implementing flash attention addresses the memory usage challenges in transformers, showcasing efforts towards computational efficiency.

<mark style="color:green;">**Incorporating Rotary Position Embeddings (RoPE)**</mark>

The use of RoPE signifies an innovative approach to incorporating positional information, crucial for understanding language sequence and structure.

<mark style="color:green;">**Special Tokens in Training for Contextual Separation**</mark>

Using end-of-text tokens to demarcate files in training data helps the model understand the boundaries of code snippets, aiding in learning and generalization.

<mark style="color:green;">**Sequence Length and Tokenization in Training**</mark>

Understanding the importance of sequence length and the tokenization process is key to how language models process and interpret code.

<mark style="color:green;">**Training with Reduced Precision for Efficiency**</mark>

Using FP16 and BFP16 in training exemplifies strategies to lessen computational load and memory requirements, reflecting efforts to make AI training more accessible.

<mark style="color:green;">**Optimising Batch Size and Learning Rate**</mark>

The choices of batch size and learning rates during training balance speed, accuracy, and overfitting risks, crucial for optimal model training.

<mark style="color:green;">**Checkpointing Strategy for Model Optimization**</mark>

Employing checkpoints during training allows for the selection of the best-performing model version, acknowledging that later training stages don't always yield better results.

<mark style="color:green;">**Hyperparameter Distinctions in Training Phases**</mark>

Differentiating hyperparameters between pretraining and fine-tuning phases is vital for tailoring the model to specific tasks without compromising its general knowledge.

<mark style="color:green;">**Enhancing API Usage Through Fine-Tuning**</mark>

Fine-tuning improves the model's proficiency in using APIs correctly, a crucial aspect of AI's practical application in software development.

<mark style="color:green;">**Challenges in API Understanding and Usage**</mark>

The limitations of language models in correctly interpreting and using evolving APIs is a significant challenge in the dynamic field of technology.

<mark style="color:green;">**Contamination in AI Benchmarking**</mark>

The issue of contamination in AI benchmarking, where training datasets might include benchmark data, underscores the need for unbiased evaluation methods.

<mark style="color:green;">**Domain**</mark> <mark style="color:green;">**Randomization for Textual Data**</mark>

Suggesting domain randomization for textual data, involving synonym replacements or altered sentence structures, aims to improve language model robustness.

<mark style="color:green;">**Customizing AI for Specific User Groups**</mark>

The concept of tailoring language models to specific regions, cultures, or demographics suggests a future of AI customization to meet diverse user needs.


# Phi 2.0

Microsoft's small but powerful transformer model

Phi-2.0, a large language model by Microsoft, can be used effectively for small, lightweight use cases in LLM applications, and these models can cooperate to create applications for business and consumers.&#x20;

### <mark style="color:purple;">Phi-2.0 Overview</mark>

* **Improved Version**: Phi-2.0 is an advancement over Phi-1.5, with doubled parameters (2.7 billion) and extended training data, making it outperform its predecessor and other larger models on several benchmarks.
* **Architecture**: It's a Transformer-based causal model, using a mix of synthetic data created with GPT-3.5 and filtered web data for training.
* **Training**: The model underwent training with 1.4 trillion tokens over 14 days using 96 A100 GPUs, displaying improved behavior in areas like toxicity and bias.

### <mark style="color:purple;">Utilisation in Lightweight Applications</mark>

* **Hardware Requirements**: Phi-2.0 is suitable for smaller setups, requiring at least 5.4 GB of GPU VRAM for fp16 parameters, and can be optimized to run on GPUs with lower VRAM by quantizing to 4-bit.
* **Fine-Tuning**: The model is easier and cheaper to fine-tune than its predecessors.&#x20;

### <mark style="color:purple;">Cooperative Application Development</mark>

* **Fine-tuning for Instructions**: Phi-2.0 can be further enhanced by fine-tuning it on instruction datasets, making it more effective in following instructions.
* **Application in Business and Consumer Products**: By leveraging its ability to be fine-tuned on smaller hardware and its improved handling of instructions, Phi-2.0 can be integrated into various business and consumer applications. These might include real-time data processing, automated customer service, language translation, content generation, and more.

### <mark style="color:purple;">Technical Implementation</mark>

* **Inference Performance**: The model shows robust inference performance with different configurations, including fp16 and 4-bit quantized versions.
* **Memory and Speed**: The quantized version of the model consumes less VRAM but at a slightly reduced inference speed. Adjustments like `flash_attn`, `flash_rotary`, and `fused_dense` can further optimize performance, especially on recent GPUs.

### <mark style="color:purple;">Conclusion</mark>

Phi-2.0's smaller size, coupled with its ability to be fine-tuned on less powerful hardware, makes it an attractive option for developing lightweight LLM applications. *<mark style="color:yellow;">**Its efficiency in handling instructions and real-time data can be particularly beneficial in creating cooperative applications for both business and consumer use**</mark>*.&#x20;


# Phi-3 Technical Report

A Highly Capable Language Model Locally on Your Phone

This <mark style="color:blue;">**April 2024**</mark> paper from the team at Microsoft introduces phi-3-mini, a 3.8 billion parameter language model that achieves performance rivalling much larger models like Mixtral 8x7B and GPT-3.5, despite being small enough to run on a phone.&#x20;

This feat was achieved solely by improving the training data, not by increasing model size.

This continues to Microsoft's work in training smaller language models such at Phi 1.5 and Phi 2.0.

Phi-3-mini's small size (3.8 billion parameters) *<mark style="color:yellow;">**enables it to run on devices like smartphones**</mark>*, opening up possibilities for local, private, and efficient applications.

{% embed url="<https://arxiv.org/abs/2404.14219>" %}
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
{% endembed %}

### <mark style="color:purple;">Key points</mark>

#### <mark style="color:green;">Training data</mark>

The *<mark style="color:yellow;">innovation lies in the dataset</mark>* used for training phi-3-mini, which is a scaled-up version of the one used for phi-2.  It consists of <mark style="color:yellow;">**heavily filtered web data and synthetic data**</mark> generated by language models.

#### <mark style="color:green;">Model architecture</mark>

phi-3-mini is a transformer decoder with a <mark style="color:yellow;">**default context length of 4K tokens**</mark>.&#x20;

It has a similar block structure to Llama-2 and uses the same tokenizer with a vocabulary size of 320,641. The model has 3072 hidden dimensions, 32 heads, and 32 layers.

#### <mark style="color:green;">Training methodology</mark>

The authors focused on the quality of data for a given scale, aiming to calibrate the training data to be closer to the "data optimal" regime for small models.&#x20;

They filtered web data to contain the correct level of "knowledge" and prioritised web pages that could potentially improve the model's reasoning ability.

#### <mark style="color:green;">Scaling results</mark>

The authors also provided initial parameter-scaling results with 7B and 14B models (phi-3-small and phi-3-medium) trained on 4.8T tokens.  These models significantly outperform phi-3-mini on benchmarks like MMLU and MT-bench.

#### <mark style="color:green;">Post-training</mark>

phi-3-mini underwent <mark style="color:blue;">**supervised fine-tuning (SFT)**</mark> and [<mark style="color:blue;">**direct preference optimization (DPO)**</mark> ](/training/the-fine-tuning-process/training-processes/direct-preference-optimization-your-language-model-is-secretly-a-reward-model)to improve its performance in math, coding, reasoning, robustness, and safety.   This process also transformed the language model into an AI assistant suitable for user interaction.

#### <mark style="color:green;">Long context version</mark>

A long context version of phi-3-mini (phi-3-mini-128K) was developed using [<mark style="color:blue;">**LongRope**</mark>](/training/the-fine-tuning-process/training-processes/longrope), *<mark style="color:yellow;">**extending the context length to 128K tokens**</mark>* while maintaining performance on par with the 4K version.

The achievement of creating a highly capable language model that can run on a phone is surprising because it challenges the assumption that larger models are always better.&#x20;

By focusing on data quality and optimizing the training process, the researchers have demonstrated that smaller models can achieve impressive results when trained on the right data.

### <mark style="color:purple;">Academic Benchmarks</mark>

The next section of the paper discusses the performance of phi-3-mini on various academic benchmarks and compares it with other models such as phi-2, Mistral-7b, Mixtral-8x7b, Gemma 7B, Llama-3-instruct-8b, and GPT-3.5.

<mark style="color:green;">**General Observation:**</mark> We want to express our doubts about the use of these academic benchmarks to assess model quality. We may be seeing situations where model developers are using techniques to improve model performance on benchmarks by including them in the training data.

### <mark style="color:blue;">Key points</mark>

#### <mark style="color:green;">Benchmarks</mark>

The models are evaluated on a wide range of tasks, including common sense reasoning (e.g., PIQA, SociQA), logical reasoning (e.g., ANLI, GSM-8K), and domain-specific knowledge (e.g., MedQA, TriviaQA). The evaluation uses few-shot prompts (varying from 0 to 10 shots) at temperature 0.

#### <mark style="color:green;">Performance</mark>

phi-3-mini (3.8B parameters) *<mark style="color:yellow;">**achieves impressive results across most benchmarks**</mark>*, often outperforming larger models like Mistral-7b and Llama-3-instruct-8b.

It even rivals the performance of GPT-3.5 on some tasks (e.g., MMLU, HellaSwag, ANLI).

#### <mark style="color:green;">Scaling results</mark>

The preview results for phi-3-small (7B) and phi-3-medium (14B) show further improvements in performance, with phi-3-medium achieving an average score of 78.2% across the benchmarks, surpassing GPT-3.5's average of 75.3%.

#### <mark style="color:green;">Coding benchmarks</mark>

phi-3-mini performs exceptionally well on coding tasks like HumanEval (59.1%) and MBPP (70.0%), outperforming larger models like Mixtral-8x7b and GPT-3.5.

Overall, the paper demonstrates that phi-3-mini achieves remarkable performance on a wide range of benchmarks while maintaining a strong focus on safety and responsible AI principles.&#x20;

The model's ability to rival larger models in terms of both performance and safety is a testament to the effectiveness of the training methodology and data optimisation techniques employed by the researchers.

### <mark style="color:purple;">Weaknesses analysis</mark>

The main weaknesses of phi-3-mini are:

<mark style="color:green;">**Limited capacity for storing factual knowledge**</mark> due to its small size, resulting in lower performance on tasks like TriviaQA that require a vast amount of factual information.

<mark style="color:green;">**Restricted language capabilities**</mark>, as the model is mostly trained on English data, limiting its multilingual performance.

<mark style="color:green;">**Challenges common to most LMs**</mark>, such as factual inaccuracies (hallucinations), reproduction or amplification of biases, inappropriate content generation, and safety issues, despite the efforts made to mitigate these problems.

### <mark style="color:blue;">Conclusion</mark>

phi-3-mini is a ground-breaking language model that demonstrates the potential of optimising training data and methodology to achieve impressive performance in a compact model size.&#x20;

Despite its limitations in storing factual knowledge and multilingual capabilities, phi-3-mini rivals the performance of much larger models on a wide range of benchmarks while maintaining a strong focus on safety and responsible AI principles.&#x20;

The model's ability to run on a phone while delivering high-quality results opens up new possibilities for applications that require on-device processing and privacy.

### <mark style="color:purple;">Three applications</mark>

#### <mark style="color:green;">**Personal AI assistant**</mark>

phi-3-mini's small size and strong performance make it an ideal candidate for a personal AI assistant that can run on a smartphone.&#x20;

Users can interact with the model directly on their devices without the need for internet connectivity, ensuring privacy and faster response times. The assistant can help with tasks such as answering questions, providing recommendations, and offering creative writing suggestions.

#### <mark style="color:green;">Educational tool</mark>

phi-3-mini's strong performance on coding tasks like HumanEval and MBPP suggests that it can be used as an educational tool for students learning programming.&#x20;

The model can provide explanations, generate code snippets, and offer guidance on coding best practices. Its ability to run on a phone makes it accessible to students in regions with limited internet access.

#### <mark style="color:green;">On-device customer support</mark>

phi-3-mini can be integrated into customer support applications that run on smartphones, allowing users to receive instant assistance *<mark style="color:yellow;">**without the need for an internet connection.**</mark>*&#x20;

The model can answer common queries, provide troubleshooting steps, and guide users through various processes. Its strong language understanding and reasoning abilities ensure that users receive accurate and helpful responses, improving customer satisfaction and reducing the workload on human support staff.

### <mark style="color:purple;">Paper References</mark>

The phi-3-mini paper references a diverse set of prior work, including research on language model scaling, training methodologies, benchmarking, and responsible AI. The key areas covered by the referenced papers are:

#### <mark style="color:green;">Language model scaling and training</mark>

* Scaling laws for neural language models \[KMH+20]
* Training compute-optimal large language models \[HBM+22]
* Scaling data-constrained language models \[MRB+23]

#### <mark style="color:green;">Transformer architecture and attention mechanisms</mark>

* Attention is all you need \[VSP+17]
* LongRope: Extending LLM context window beyond 2 million tokens \[DZZ+24]

#### <mark style="color:green;">Previous work on phi models and training data optimisation</mark>

* Textbooks are all you need \[GZA+23]
* Textbooks are all you need ii: phi-1.5 technical report \[LBE+23]
* phi-2: The surprising power of small language models \[JBA+23]

#### <mark style="color:green;">Benchmarking and evaluation</mark>

* MMLU \[HBK+21], HellaSwag \[ZHB+19], ANLI \[NWD+20], GSM-8K \[CKB+21], MedQA \[JPO+20], AGIEval \[ZCG+23], TriviaQA \[JCWZ17], Arc-C/Arc-E \[CCE+18], PIQA/SociQA \[BZGC19], BigBench-Hard \[SRR+22, SSS+22], WinoGrande \[SLBBC19], OpenBookQA \[MCKS18], BoolQ \[CLC+19], CommonSenseQA \[THLB19], TruthfulQA \[LHE22], HumanEval \[CTJ+21], MBPP \[AON+21], GPQA \[RHS+23], MTBench \[ZCS+23]

#### <mark style="color:green;">Responsible AI and safety alignment</mark>

* Training a helpful and harmless assistant with reinforcement learning from human feedback \[BJN+22]
* Beavertails: Towards improved safety alignment of LLM via a human-preference dataset \[JLD+23]
* Safety-tuned LLaMas: Lessons from improving the safety of large language models that follow instructions \[BSA+24]

#### <mark style="color:green;">Other related language models</mark>

* GPT-2 \[RWC+19], Llama \[TLI+23], Mistral \[JSM+23], Mixtral \[JSR+24], Gemma \[TMH+24]


# The Fine Tuning Process

Fine tuning deep learning models is completely different to fine tuning machine learning models

***

\
Fine-tuning is a process that enhances the performance of neural language models by precisely adjusting them for specific tasks.&#x20;

Unlike prompt engineering, which directs a model without altering its internal structure, fine-tuning involves a deeper modification, adjusting the model's internal parameters, such as weights, to improve its task-specific performance.&#x20;

<figure><img src="/files/8JyFWZoNWBRmdVP83Iw0" alt=""><figcaption></figcaption></figure>

This process not only sharpens the model's abilities, ensuring higher accuracy and reliability but also saves significant time and resources by leveraging pre-existing knowledge from pre-trained models.

Fine-tuning equips models with the flexibility to adapt to a wide range of tasks with minimal adjustments, making it an essential step in the model training pipeline.

After pretraining, fine-tuning refines the broad knowledge base of the model, aligning it with the particularities of a new task, which significantly optimises performance.&#x20;

The process involves preparing diverse and relevant data, updating model weights through backpropagation, and carefully selecting hyperparameters to guide the model's learning.&#x20;

Fine-tuning is not just an adjustment but an enhancement that ensures neural language models achieve task-specific precision and efficiency.&#x20;


# Why fine tune?

There is an ongoing debate around the need for fine tuning a large language model.  Many suggest they have found no need for it, that prompt engineering and 'retrieval augmented generation' suffices.

Given fine tuning is difficult and resource intensive, it is not surprising it can be ignored.  But it should not be.  It is a powerful tool in the kit to augment an AI application.

Generally the debate is fine tuning versus Retrieval Augmented Generation (RAG).  This debate is often framed as an either/or choice, but in reality, these techniques serve complementary functions that can be synergistically integrated for optimal results.

In this section of our documentation, we will review the academic work and other sources to make the case for fine tuning and when and how it should be used.

### <mark style="color:purple;">Fine-Tuning</mark>&#x20;

**Fine-tuning** involves adjusting the weights of a pre-trained language model to improve its performance on specific tasks. This is akin to specialty training in various professions—just as a doctor undergoes specialty training to make precise diagnoses, a language model may be fine-tuned with domain-specific data to enhance its task-specific accuracy. For instance, a model could be fine-tuned on legal texts to better perform tasks related to legal analysis.

**Examples of Fine-Tuning:**

* **Language Modeling Task Fine-tuning**: Adapts a pre-trained model to improve its general linguistic capabilities or to refine its skills in generating coherent and contextually appropriate text.
* **Supervised Q\&A Fine-tuning**: Specifically enhances the model's abilities in question-answering scenarios by training it on a dataset of question and answer pairs.

### <mark style="color:purple;">Retrieval Augmented Generation (RAG)</mark>

**Retrieval Augmented Generation** enhances a model's responses by integrating external data into the model's context at inference time. This method can be compared to a doctor consulting a patient's medical history before making a diagnosis. RAG allows a model to access a wide array of information that isn't stored in its parameters but can be crucial for generating accurate and informed outputs.

**Examples of RAG:**

* **Using vector databases** to pull in relevant information based on the query context.
* **Incorporating data from APIs or traditional databases** to provide real-time, relevant information that the model can use to generate responses.

### <mark style="color:purple;">Combining FT and RAG</mark>

Integrating FT and RAG can significantly enhance a model's performance by not only refining its internal understanding and response generation but also by expanding its access to and use of external information.&#x20;

For example, in a healthcare application, a model might be fine-tuned with medical research to understand and generate medically accurate text while also using RAG to pull patient-specific information from medical records to tailor its responses to individual cases.

**Example of Combined Use:**

* A customer service LLM could be fine-tuned on high-quality customer interaction logs to learn the best communicative practices while using RAG to pull user-specific data to personalise interactions, such as recommending products based on past purchases or addressing past complaints.

### <mark style="color:purple;">Conclusion</mark>

The choice between fine-tuning and RAG should not be seen as a binary one; each has its strengths and applications.&#x20;

Fine-tuning allows for deep customisation of the model's behavior and understanding, making it more adept at specific tasks.&#x20;

RAG, on the other hand, supplements the model's capabilities by providing additional, context-relevant information at runtime, making it adaptable and resourceful.&#x20;

When used together, they provide a robust framework for developing powerful, context-aware, and highly specialised LLM applications.&#x20;

This approach empowers developers to leverage the strengths of both techniques to build more dynamic, responsive, and effective models.

4


# Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?

In this <mark style="color:blue;">**May 2024**</mark> paper the authors explore the impact of fine-tuning large language models (LLMs) on new, previously unlearned factual information.&#x20;

The study focuses on the hypothesis that exposure to such new knowledge during fine-tuning may increase the likelihood of the models generating factually incorrect responses, a phenomenon known as hallucination.&#x20;

Using a controlled setup with closed-book question answering, the authors vary the proportion of fine-tuning examples that introduce new knowledge and observe the models' performance.&#x20;

Their findings reveal that while LLMs struggle to learn new factual information through fine-tuning, eventually incorporating this new knowledge increases the models' propensity to hallucinate.&#x20;

These results underscore the <mark style="color:yellow;">potential risks associated with introducing new knowledge via fine-tuning</mark> and suggest that LLMs primarily acquire factual knowledge during pre-training, with fine-tuning enhancing their ability to use this knowledge effectively.

{% embed url="<https://arxiv.org/abs/2405.05904>" %}
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?" by Zorik Gekhman et al.
{% endembed %}

### <mark style="color:purple;">Why should you bother with fine-tuning?</mark>&#x20;

Fine-tuning aligns LLMs with desired behaviours and adapting them to specific downstream tasks.&#x20;

It allows you to leverage the general knowledge acquired by LLMs during pre-training and tailor it to your specific use case.&#x20;

Fine-tuning can significantly improve the performance and utility of LLMs for practical applications.

#### <mark style="color:blue;">Benefits of fine-tuning</mark>

1. Improved performance on specific tasks compared to using the pre-trained LLM directly.
2. Ability to adapt the LLM to domain-specific language, terminology, and style.
3. Opportunity to teach the LLM to follow instructions and exhibit desired behaviours.
4. Potential to enhance the LLM's capability to utilize its pre-existing knowledge effectively.

#### <mark style="color:blue;">When to fine-tune</mark>

1. When you have a specific downstream task or application that requires the LLM to follow certain instructions or exhibit specific behaviours.
2. When you need the LLM to adapt to domain-specific language, terminology, or style.
3. When you want to improve the LLM's performance on a particular task or set of tasks relevant to your use case.

#### <mark style="color:blue;">Best practices for fine-tuning</mark>

1. Use high-quality, task-specific data for fine-tuning that aligns with the desired behavior and domain.
2. Be <mark style="color:yellow;">cautious about introducing new factual knowledge</mark> through fine-tuning data, as it may encourage hallucinations. Consider filtering out or re-labelling examples that introduce new facts.
3. Employ early stopping based on a validation set to mitigate overfitting and reduce the risk of hallucinations.
4. Carefully select the fine-tuning examples to include a mix of HighlyKnown and MaybeKnown examples, as they are essential for the LLM to use its pre-existing knowledge effectively.

<mark style="color:blue;">**Note:**</mark> The paper provides evidence that fine-tuning works well when the fine-tuning dataset consists *<mark style="color:yellow;">primarily of examples that are known to the pre-trained LLM</mark>* (referred to as "Known" examples in the paper).  The authors demonstrate that fine-tuning on a dataset with a higher proportion of "Known" examples leads to better performance on a held-out test set.  Conversely, fine-tuning on a dataset with a higher proportion of examples containing new knowledge that the LLM was not exposed to during pre-training (referred to as "Unknown" examples) results in decreased performance and a higher tendency for the model to hallucinate.

The authors demonstrated this by categorising the "Known" examples into three subcategories: HighlyKnown, MaybeKnown, and WeaklyKnown.&#x20;

They show that fine-tuning on a dataset consisting solely of HighlyKnown examples leads to suboptimal performance, as the model struggles to handle MaybeKnown examples during inference. On the other hand, fine-tuning on a dataset with a mix of HighlyKnown and MaybeKnown examples results in the best overall performance, as it allows the LLM to effectively use its pre-existing knowledge across all subcategories of Known examples.

### <mark style="color:purple;">Practical and commercial applications of fine-tuning</mark>

Fine-tuning enables LLMs to adapt to specific tasks and domains by leveraging the knowledge acquired during pre-training while learning to apply it in a targeted manner.&#x20;

#### <mark style="color:blue;">Developing domain-specific chatbots or virtual assistants</mark>

Fine-tuning allows LLMs to learn the language, terminology, and common queries specific to a particular domain, such as customer support or sales. By training on domain-specific data, the LLM can generate more relevant and accurate responses, leading to improved user experience and satisfaction.

<mark style="color:blue;">Creating specialised content generation tools</mark>

Fine-tuning enables LLMs to learn the style, tone, and structure of content specific to a domain, such as marketing copy, news articles, or creative writing. By exposing the LLM to high-quality examples during fine-tuning, it can generate content that closely mimics the desired style and meets the specific requirements of the target domain.

<mark style="color:blue;">Building knowledge retrieval systems</mark>

Fine-tuning can teach LLMs to identify and retrieve relevant information from a domain-specific knowledge base. By training on examples of questions and their corresponding answers, the LLM learns to understand the context and intent behind user queries and provide accurate and concise responses.

#### <mark style="color:blue;">Adapting LLMs for task-specific applications</mark>

Fine-tuning allows LLMs to specialise in tasks like summarisation, translation, or sentiment analysis by learning from task-specific training data.&#x20;

For example, fine-tuning on a dataset of document-summary pairs teaches the LLM to identify key information and generate coherent summaries, while fine-tuning on a dataset of text-sentiment pairs enables the LLM to accurately classify the sentiment expressed in a given piece of text.

<mark style="color:blue;">Fine-tuning LLMs for educational purposes</mark>

Fine-tuning can adapt LLMs to generate educational content, such as explanations, quizzes, or personalised learning materials.  By training on a dataset of educational content and student interactions, the LLM can learn to generate content that is tailored to the learner's needs, level of understanding, and learning style, ultimately improving the learning experience and outcomes.

In summary, fine-tuning enables LLMs to acquire domain-specific knowledge, learn task-specific patterns and structures, and generate outputs that closely align with the desired behavior and objectives. This adaptability and specialization make fine-tuned LLMs valuable tools for a wide range of practical and commercial applications.


# Explanations in Fine Tuning

This <mark style="color:blue;">February 2024</mark> paper suggest that the inclusion of explanations can enable models to solve complex problem-solving tasks more effectively than traditional training methods.&#x20;

The process of fine-tuning can be complex and resource-intensive, often requiring large amounts of data and computational power.&#x20;

This study has shed light on how the inclusion of *<mark style="color:yellow;">explanations in the training data can significantly enhance the fine-tuning process</mark>*, leading to improved performance and more efficient learning.

The research team's findings demonstrate that by incorporating step-by-step explanations into the training data, language models can achieve higher accuracy, solve previously unsolvable tasks, and generalize better to new challenges.

{% embed url="<https://arxiv.org/abs/2402.07543>" %}

### <mark style="color:purple;">The key findings</mark>

The inclusion of explanations within the training data significantly boosts the performance of language models, particularly helping smaller models to a greater extent than larger ones.

<mark style="color:blue;">**Evidence:**</mark> The T5-small model (60 million parameters) achieved 87.8% accuracy with long explanations, compared to 65.1% without explanations. Larger models like T5-3B (2.7 billion parameters) also benefited from explanations but to a lesser degree, achieving 99.3% accuracy with long explanations compared to 65.8% without.

Models fine-tuned with explanations can solve tasks they previously could not handle, indicating that explanations help bridge gaps in a model's knowledge and reasoning capabilities.

<mark style="color:blue;">**Evidence:**</mark> On the modular sum task, all models performed no better than random guessing without explanations. However, with explanations, the models could solve the problem, with T5-small achieving 75.7% accuracy with long explanations and larger models scoring over 98% with any type of explanation.

Adding explanations to the training data not only reduces the required volume of data but also facilitates better generalization across tasks.

<mark style="color:blue;">**Evidence:**</mark> The T5-large model trained on just 2,000 explained samples achieved 66.1% accuracy, outperforming the model trained on 2,000 unexplained samples (44.9%). When trained on sequences of length 50-100 and tested on sequences of length 100-200, the T5-large model achieved 91.2% accuracy with medium-length explanations, compared to 63.5% without explanations.

The complexity of explanations impacts the model's learning curve over time, with diminishing returns on model performance improvements with increasingly detailed explanations.

<mark style="color:blue;">**Evidence:**</mark> For the T5-small model, longer explanations led to faster convergence and higher accuracy. However, for larger models like T5-base and T5-large, the effect of explanation length on convergence speed and final accuracy was less pronounced, suggesting that larger models require less detailed explanations to benefit.

In summary, this paper provides strong evidence for the benefits of fine-tuning language models with explanations, particularly for smaller models and complex problem-solving tasks.&#x20;

The inclusion of explanations can reduce the required training data, improve generalization, and help models solve tasks they previously could not handle.

However, the optimal level of explanation complexity may vary depending on the model size, and there may be limitations to the model's ability to generalise to certain types of variations in the input data.

<figure><img src="/files/3co1rfvwU1UirATxS1LM" alt="" width="549"><figcaption><p>Shows the results of the T5-small and T5-3B models on ListOps, comparing performances when fine-tuned without explanations versus with three the types of explanations.</p></figcaption></figure>

### <mark style="color:purple;">Instruction Tuning Dataset: Movie Genre Classification</mark>

### <mark style="color:blue;">No Explanations</mark>

{% code overflow="wrap" %}

```yaml
Instruction: Classify the following movie plot summary into one of three genres: action, comedy, or drama.

Input: When a retired CIA agent is kidnapped by a terrorist group, his son, a young MIT graduate, takes matters into his own hands and embarks on a daring rescue mission.

Output: action
```

{% endcode %}

{% code overflow="wrap" %}

```yaml
Instruction: Classify the following movie plot summary into one of three genres: action, comedy, or drama.

Input: A quirky, socially awkward woman moves to a small town and becomes entangled in the lives of her eccentric new neighbors, leading to hilarious misunderstandings and unexpected friendships.

Output: comedy
```

{% endcode %}

{% code overflow="wrap" %}

```yaml
Instruction: Classify the following movie plot summary into one of three genres: action, comedy, or drama.

Input: In the wake of a tragic accident, a once-promising athlete must confront his inner demons and rebuild his life while navigating complex relationships with his family and friends.

Output: drama
```

{% endcode %}

### <mark style="color:blue;">Short Explanations</mark>

{% code overflow="wrap" %}

```yaml
Instruction: Classify the following movie plot summary into one of three genres and briefly explain your choice: action, comedy, or drama.

Input: When a retired CIA agent is kidnapped by a terrorist group, his son, a young MIT graduate, takes matters into his own hands and embarks on a daring rescue mission.

Output: action. The plot involves a rescue mission and a CIA agent, which are common elements in action movies.
```

{% endcode %}

{% code overflow="wrap" %}

```yaml
Instruction: Classify the following movie plot summary into one of three genres and briefly explain your choice: action, comedy, or drama.

Input: A quirky, socially awkward woman moves to a small town and becomes entangled in the lives of her eccentric new neighbors, leading to hilarious misunderstandings and unexpected friendships.

Output: comedy. The plot features quirky characters, misunderstandings, and humor, which are typical of comedy movies.
```

{% endcode %}

{% code overflow="wrap" %}

```yaml
Instruction: Classify the following movie plot summary into one of three genres and briefly explain your choice: action, comedy, or drama.

Input: In the wake of a tragic accident, a once-promising athlete must confront his inner demons and rebuild his life while navigating complex relationships with his family and friends.

Output: drama. The plot deals with a character's inner struggles and complex relationships, which are often the focus of drama movies.
```

{% endcode %}

### <mark style="color:blue;">Medium Explanations</mark>

{% code overflow="wrap" %}

```yaml
Instruction: Classify the following movie plot summary into one of three genres and provide a medium-length explanation for your choice: action, comedy, or drama.

Input: When a retired CIA agent is kidnapped by a terrorist group, his son, a young MIT graduate, takes matters into his own hands and embarks on a daring rescue mission.

Output: action
- The plot involves a kidnapping by a terrorist group, which creates a high-stakes situation.
- The main character is a young MIT graduate, suggesting intelligence and resourcefulness.
- The son embarks on a daring rescue mission, implying action and adventure.
Based on these elements, the movie is likely an action film.
```

{% endcode %}

{% code overflow="wrap" %}

```yaml
Instruction: Classify the following movie plot summary into one of three genres and provide a medium-length explanation for your choice: action, comedy, or drama.

Input: A quirky, socially awkward woman moves to a small town and becomes entangled in the lives of her eccentric new neighbors, leading to hilarious misunderstandings and unexpected friendships.

Output: comedy
- The main character is described as quirky and socially awkward, which often leads to humorous situations.
- The plot involves misunderstandings and unexpected friendships, suggesting a lighthearted tone.
- The small-town setting and eccentric neighbors are common tropes in comedy movies.
Given these factors, the movie is most likely a comedy.
```

{% endcode %}

{% code overflow="wrap" %}

```yaml
Instruction: Classify the following movie plot summary into one of three genres and provide a medium-length explanation for your choice: action, comedy, or drama.

Input: In the wake of a tragic accident, a once-promising athlete must confront his inner demons and rebuild his life while navigating complex relationships with his family and friends.

Output: drama
- The plot centers around a tragic accident, which sets a serious and emotional tone.
- The main character must confront his inner demons, indicating a focus on personal struggles and growth.
- The character navigates complex relationships with family and friends, suggesting interpersonal drama.
Considering these elements, the movie falls into the drama genre.
```

{% endcode %}

### <mark style="color:blue;">Long Explanations</mark>

{% code overflow="wrap" %}

```yaml
Instruction: Classify the following movie plot summary into one of three genres and provide a detailed explanation for your choice: action, comedy, or drama.

Input: When a retired CIA agent is kidnapped by a terrorist group, his son, a young MIT graduate, takes matters into his own hands and embarks on a daring rescue mission.

Output: action
This movie plot summary contains several key elements that strongly suggest it belongs to the action genre:
1. The inciting incident involves the kidnapping of a retired CIA agent by a terrorist group. This sets up a high-stakes, dangerous situation that is common in action movies.
2. The protagonist is the son of the kidnapped CIA agent, and he is described as a young MIT graduate. This implies that he is intelligent, resourceful, and likely possesses skills that will aid him in the rescue mission.
3. The son "takes matters into his own hands" and embarks on a "daring rescue mission." This indicates that the plot will involve thrilling action sequences, chases, and potentially violent confrontations as the son attempts to save his father.
4. The involvement of a terrorist group suggests that there may be broader political or societal implications at play, which is another common theme in action movies.
Overall, the combination of a high-stakes kidnapping, a resourceful protagonist, and a daring rescue mission firmly places this movie in the action genre.
```

{% endcode %}

{% code overflow="wrap" %}

```yaml
Instruction: Classify the following movie plot summary into one of three genres and provide a detailed explanation for your choice: action, comedy, or drama.

Input: A quirky, socially awkward woman moves to a small town and becomes entangled in the lives of her eccentric new neighbors, leading to hilarious misunderstandings and unexpected friendships.

Output: comedy
This movie plot summary exhibits several characteristics that are strongly associated with the comedy genre:
1. The protagonist is described as "quirky" and "socially awkward." These traits often serve as a source of humor in comedies, as the character's unusual behavior or perspective can lead to amusing situations and misunderstandings.
2. The setting is a small town, which is a common trope in comedy movies. Small towns often have a close-knit community and a cast of colorful, eccentric characters that can provide a rich source of comedic material.
3. The plot involves the protagonist becoming "entangled in the lives of her eccentric new neighbors." This suggests that there will be a series of humorous interactions, misunderstandings, and potentially embarrassing situations as the character navigates her new social environment.
4. The phrase "leading to hilarious misunderstandings" directly indicates that the movie will feature plenty of comedic moments arising from the character's interactions with her neighbors.
5. The mention of "unexpected friendships" suggests that the movie will have a heartwarming aspect as well, which is common in many comedies. The protagonist's journey may involve personal growth and the formation of unlikely bonds with her quirky neighbors.
In summary, the quirky protagonist, small-town setting, eccentric neighbors, and emphasis on hilarious misunderstandings and unexpected friendships all point to this movie being a clear example of the comedy genre.
```

{% endcode %}

{% code overflow="wrap" %}

```yaml
Instruction: Classify the following movie plot summary into one of three genres and provide a detailed explanation for your choice: action, comedy, or drama.

Input: In the wake of a tragic accident, a once-promising athlete must confront his inner demons and rebuild his life while navigating complex relationships with his family and friends.

Output: drama
This movie plot summary contains several elements that strongly indicate it belongs to the drama genre:
1. The inciting incident is a "tragic accident," which immediately sets a serious and emotionally heavy tone. Dramas often explore characters grappling with difficult, life-altering events and their aftermath.
2. The protagonist is described as a "once-promising athlete," suggesting that the accident has derailed their life and career. This sets up a narrative of personal struggle and potential redemption, which is a common theme in drama movies.
3. The phrase "confront his inner demons" indicates that the movie will delve into the protagonist's psychological and emotional struggles. Dramas often focus on characters' inner lives, their flaws, and their efforts to overcome personal challenges.
4. The protagonist must "rebuild his life," which implies a journey of growth, healing, and self-discovery. This type of character arc is a hallmark of the drama genre.
5. The movie also involves the protagonist "navigating complex relationships with his family and friends." Dramas often explore interpersonal dynamics, family issues, and the ways in which relationships can be strained or strengthened during difficult times.
6. The words "tragic," "inner demons," and "complex relationships" all suggest that the movie will have a serious, emotionally intense tone, which is characteristic of the drama genre.
Overall, the combination of a tragic event, a protagonist confronting personal struggles, a journey of rebuilding and self-discovery, and complex interpersonal relationships firmly situates this movie in the drama genre.
```

{% endcode %}

This dataset provides instructions, inputs, and outputs for the task of movie genre classification, with varying levels of explanation.&#x20;

By training a language model on this dataset, it can learn to classify movie plot summaries into genres and provide explanations for its choices, depending on the level of detail requested in the instruction.

### <mark style="color:purple;">References</mark>

The references can be categorised into logical groups based on their primary focus and contributions to the field of language model research:

#### 1. **Language Model Enhancements and Applications**

* **Prompting and Fine-Tuning Techniques**: Papers discussing innovative techniques to enhance model performance through prompting or fine-tuning strategies. This includes works by Wei et al. (2022b) on "Chain-of-Thought" prompting and Ziegler et al. (2019) on fine-tuning models based on human preferences.
* **Transformer Architectures and Applications**: Seminal works on transformer architectures such as Vaswani et al. (2017), and their applications to various tasks, such as Pegasus by Zhang et al. (2020) for summarization.

#### 2. **Model Explanation and Interpretability**

* **Explanations in Machine Learning**: Papers focused on enhancing understanding of model decisions, such as Camburu et al. (2018) with e-SNLI and Hase et al. (2020) discussing the roles of explanations in model training.
* **Analyzing Model Behavior**: Studies like Ballout et al. (2023a) that explore the internal mechanisms of models, such as attention weights, for better interpretability.

#### 3. **Generalization and Multi-task Learning**

* **Cross-Domain and Multi-task Learning**: Papers examining the capabilities of language models across different tasks and domains, such as the work by Ballout et al. (2023b) on cross-domain datasets and Lu et al. (2021) on using pre-trained transformers as universal computation engines.
* **Meta-Learning and Few-Shot Learning**: Insights from Chen et al. (2022) and Brown et al. (2020) on how language models can adapt to new tasks with minimal examples.

#### 4. **Methodological Innovations in Training Language Models**

* **Training and Scaling Models**: Works that focus on novel training methods or scaling up models, such as Cobbe et al. (2021) on training verifiers and Chung et al. (2022) on scaling instruction-tuned language models.
* **Fine-Tuning and Instruction Tuning**: Studies like Liu et al. (2022) that compare different fine-tuning methods with in-context learning for efficiency and efficacy.

#### 5. **Model Reasoning and Decision Making**

* **Advanced Reasoning Strategies**: Research on advanced model reasoning techniques, such as the "Tree of Thoughts" method by Yao et al. (2023) and multimodal reasoning as explored by Zhang et al. (2023).
* **Natural Language Understanding and Reasoning**: Contributions to understanding and enhancing reasoning in language models, including Rajani et al. (2019) on leveraging language models for commonsense reasoning.

These categories reflect the diverse approaches and methodologies currently being explored in the field of language modeling, each contributing to the overarching goal of enhancing model performance, understanding, and utility across a range of applications.


# Tokenization

Tokenization is a fundamental concept in the training of large language models. &#x20;

The *<mark style="color:yellow;">**process breaks down text into smaller, manageable units called tokens**</mark>*. These tokens, which can range from individual characters to entire words, enable neural models to better understand and process human language.

Tokenization can be a complex process, as it involves handling different types of text data, such as punctuation, numbers, and special characters, and determining how to split them into meaningful units.&#x20;

Tokenization can also vary depending on the specific task or application. For example, in some cases, it may be necessary to split words into smaller subwords to handle out-of-vocabulary (OOV) words that are not present in a pre-trained vocabulary.

Once text has been tokenized, it can be further processed using techniques such as stemming, lemmatization, or part-of-speech tagging, or fed into a machine learning model for training or inference.

### <mark style="color:purple;">Tokenization Process</mark>

<mark style="color:green;">**Segmentation:**</mark> The first step of tokenization involves breaking down text into units. These units can be as large as sentences, as small as characters, or more commonly, words and subwords.

<mark style="color:green;">**Vocabulary Building:**</mark> Once you decide on the granularity of the units (words, subwords, etc.), you create a vocabulary or a list of unique tokens from the corpus.

<mark style="color:green;">**Mapping:**</mark> Each unique token in the vocabulary is assigned a unique integer ID.

<mark style="color:green;">**Encoding:**</mark> The original text is then converted or "encoded" into a sequence of these integer IDs according to the mapping. For example, if you tokenize the sentence "I love AI", and your vocabulary mapping is {'I': 1, 'love': 2, 'AI': 3}, the encoded sentence becomes \[1, 2, 3].

### <mark style="color:purple;">Here's a general outline of how to tokenize an entire dataset</mark>

<mark style="color:green;">Choose a tokenization method</mark>

Depending on the language, dataset, and the specific requirements of your task, select an appropriate tokenization method. This could be word-based, subword-based (e.g., BPE, WordPiece, or SentencePiece), or character-based tokenization.

Before tokenizing, clean and pre-process the dataset to ensure consistency and remove any irrelevant information. This might involve converting the text to lowercase, removing special characters or punctuation, or handling contractions and abbreviations.

#### <mark style="color:green;">Train the tokenizer</mark>

If you're using a data-driven tokenization method like BPE, WordPiece, or SentencePiece, you need to *train the tokenizer on your dataset*. This step allows the tokenizer to learn the most frequent and meaningful tokens in the data.

#### <mark style="color:green;">Tokenize the dataset</mark>

Once the tokenizer is trained, apply it to the entire dataset. The tokenizer will break down the text into smaller units according to the chosen method. For instance, it may convert sentences into lists of words or subword tokens.

#### <mark style="color:green;">Post-processing</mark>

After tokenization, you may want to perform additional processing steps such as adding special tokens (e.g., \[CLS], \[SEP], or \[MASK] in BERT), padding sequences to a fixed length, or creating batches for input to your NLP model.

#### <mark style="color:green;">Save the tokenized dataset</mark>

Finally, save the tokenized dataset for further use in training or evaluation tasks.&#x20;

Depending on the downstream task, you may want to *save the dataset in a specific format* (e.g., PyTorch tensors, TensorFlow tensors, or NumPy arrays).


# Tokenization Is More Than Compression

Craig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, Chris Tanner

This <mark style="color:blue;">**February 2024**</mark> paper explores the role of tokenization in Natural Language Processing (NLP) tasks and challenges the common understanding of why certain tokenization methods, such as Byte-Pair Encoding (BPE), are effective.

{% embed url="<https://arxiv.org/abs/2402.18376>" %}
Tokenization Is More Than Compression
{% endembed %}

<mark style="color:yellow;">Tokenization is the process of converting raw text into a sequence of distinct tokens that can be used by statistical models.</mark>&#x20;

The authors divide tokenization into three stages:

<mark style="color:blue;">**Pre-tokenization:**</mark> Optional initial rules that restrict or enforce the creation of certain tokens (e.g., splitting a corpus on whitespace).

<mark style="color:blue;">**Vocabulary Construction**</mark><mark style="color:blue;">:</mark> The core algorithm that constructs a vocabulary of tokens (V) of size m from a given text corpus (C), while adhering to pre-tokenization rules.

<mark style="color:blue;">**Segmentation:**</mark> The process of splitting a document (d) into a series of tokens (t1, ..., tKd) from the vocabulary (V), such that the concatenation of the tokens equals the original document.

The authors introduce a new metric called <mark style="color:blue;">**Corpus Token Count (CTC),**</mark> which is the total number of tokens used in the segmentation of all documents in a corpus.

The paper *<mark style="color:yellow;">**challenges the hypothesis that the effectiveness of BPE stems from its ability to compress text into a short sequence of tokens.**</mark>*&#x20;

To test this, they introduce a novel tokenizer called PATHPIECE, which finds a segmentation with the minimum possible number of tokens (Kd) for a given document and vocabulary.   The PATHPIECE vocabulary construction routine is a top-down procedure that heuristically minimizes CTC on a training corpus.

The authors conduct experiments by training 64 language models (LMs) with varying tokenization methods and vocabulary sizes:

* 54 LMs with 350M parameters
* 6 LMs with 1.3B parameters
* 4 LMs with 2.4B parameters

They evaluate the impact of different tokenization stages and vocabulary sizes on downstream task performance.  The paper provides open-source access to PATHPIECE, token vocabularies, and all 64 trained LMs.

### <mark style="color:purple;">Related Work</mark>

The related work section discusses pre-tokenization methods, vocabulary construction algorithms, and segmentation methods in detail.

#### <mark style="color:green;">Pre-tokenization Methods</mark>

Pre-tokenization is the process of breaking text into chunks, which are then tokenized independently.&#x20;

Tokens are not allowed to cross pre-tokenization boundaries.  The authors discuss three pre-tokenization methods:

<mark style="color:blue;">**FirstSpace:**</mark> Used by BPE, WordPiece, and Unigram, it requires new chunks to begin whenever a space is encountered. If a space appears in a chunk, it must be the first character.

<mark style="color:blue;">**Space:**</mark> Suggested by Gow-Smith et al. (2022), it treats spaces as individual tokens.

<mark style="color:blue;">**Digit:**</mark> Popularized by Llama (Touvron et al., 2023), it treats each digit as an individual token.

#### <mark style="color:green;">Vocabulary Construction</mark>&#x20;

The authors focus on byte-level, lossless subword tokenization algorithms that split text into word and subword units based on their frequency and co-occurrence patterns from their "training" data. They analyse four subword tokenizers:

<mark style="color:blue;">**Byte-Pair Encoding (BPE):**</mark> A bottom-up method that starts with single bytes as tokens and merges the most commonly occurring pair of adjacent tokens in a training corpus into a single new token until the desired vocabulary size is reached.

<mark style="color:blue;">**WordPiece:**</mark> Similar to BPE, but uses Pointwise Mutual Information (PMI) as the criteria to identify candidates to merge, prioritizing pairs that occur together more frequently than expected, relative to the individual token frequencies.

<mark style="color:blue;">**Unigram Language Model**</mark><mark style="color:blue;">:</mark> A top-down approach that starts from a large initial vocabulary and progressively prunes groups of tokens that induce the minimum likelihood decrease of the corpus, selecting tokens to maximise the likelihood of the corpus according to a simple unigram language model.

<mark style="color:blue;">**SaGe**</mark><mark style="color:blue;">:</mark> Proposed by Yehezkel and Pinter (2023), it incorporates contextual information into an ablation loss via a skip-gram objective and operates top-down, pruning from an initial vocabulary to a desired size.

### <mark style="color:green;">Segmentation Methods</mark>

Segmentation converts text into a series of tokens, given a tokenizer and a vocabulary of tokens.&#x20;

The authors ensure that all 256 single-byte tokens are included in the vocabulary to avoid out-of-vocabulary issues.&#x20;

Some segmentation methods are tightly coupled to the vocabulary construction step, such as merge rules for BPE or the maximum likelihood approach for Unigram.  Others, like the WordPiece approach of greedily taking the longest prefix token in the vocabulary at each point, can be applied to any vocabulary.&#x20;

Alternative segmentation schemes include Dynamic Programming BPE, BPE-Dropout, and FLOTA.

### <mark style="color:purple;">Experiments</mark>

The experiments section provides details on the authors' experimental setup, the downstream evaluation tasks used, and the various tokenization stage variants tested.&#x20;

The authors then present and analyse the results of their experiments.

#### <mark style="color:green;">Downstream Evaluation Tasks</mark>

The authors selected 10 benchmarks from the lm-evaluation-harness to evaluate the performance of their tokenization process.&#x20;

These benchmarks are all multiple-choice tasks with 2, 4, or 5 options and were run with 5-shot prompting.

#### <mark style="color:green;">Tokenization Stage Variants</mark>

The authors conducted 18 experimental variants, each repeated at vocabulary sizes of 32,768, 40,960, and 49,152.&#x20;

They used BPE, Unigram, WordPiece, and SaGe as baseline vocabulary creation methods, and two variants of PATHPIECE with different tie-breaking strategies (longest token and random). They also varied the initial vocabulary for PATHPIECE and SaGe, and the pre-tokenization schemes.

### <mark style="color:purple;">Results</mark>

The authors reports the downstream performance across all experimental settings. They make several key observations:

<mark style="color:blue;">**Vocabulary Size:**</mark> The authors found a high correlation between downstream performance at different vocabulary sizes, indicating that vocabulary size is not a crucial decision over the range of 30k to 50k tokens.

<mark style="color:blue;">**Overall Performance:**</mark> The top five tokenizers (PATHPIECEL-BPE, Unigram, BPE, BPE-Greedy, and WordPiece) do not have any statistically significant differences in performance. SaGe-BPE (rank 6) is only barely worse than PATHPIECEL-BPE. This suggests that there is no single tokenizer algorithm that is significantly better than the others.

<mark style="color:blue;">**Model Size:**</mark> The authors built larger models (1.3B and 2.4B parameters) for a subset of the experiments. They found that the relative performance of the tokenizers varies by model size, but there is still a group of highly-performant tokenizers that yield comparable results.

#### <mark style="color:green;">Corpus Token Count vs Accuracy</mark>&#x20;

The authors did not find a straightforward relationship between the corpus token count (CTC) versus the accuracy of each vocabulary size.&#x20;

The range of CTC is quite narrow within each vocabulary construction method, even while changes in pre-tokenization and segmentation lead to significant accuracy differences. The authors suggest that there might be an inverted U-shaped curve with respect to the CTC and downstream performance.

In summary, the experiments and results demonstrate that there is no single tokenizer algorithm that significantly outperforms the others, and the relationship between corpus token count and downstream performance is not straightforward.

&#x20;The authors' findings suggest that factors other than vocabulary size and corpus token count play a role in the effectiveness of tokenization for language modelling tasks.

### <mark style="color:purple;">Conclusion</mark>

In this paper, the authors investigate the hypothesis that reducing the corpus token count (CTC) would improve downstream performance in natural language processing tasks. &#x20;

They compare various tokenization methods and analyse the impact of different stages of tokenization on downstream task performance.

The *<mark style="color:yellow;">**main conclusion is that the relationship between CTC and downstream accuracy is not straightforward,**</mark>* as five different tokenizers with varying CTCs perform comparably.  This finding challenges the current understanding of why Byte-Pair Encoding (BPE) is particularly effective.&#x20;

The authors also find that unigram tokenizers better align with morphological segmentation compared to BPE tokenizers, further suggesting that the effectiveness of a tokenizer cannot be explained by a single factor.


# Tokenization - SentencePiece

The Unsupervised Text Tokenizer for Neural Networks

SentencePiece is a language-independent subword tokenizer and detokenizer designed for Neural-based text processing, particularly Neural Machine Translation (NMT).&#x20;

Its main goal is to provide a simple, efficient, and reproducible preprocessing and postprocessing tool that can be easily integrated into neural network-based NLP systems.

Its core strength lies in its ability to *<mark style="color:yellow;">**manage vocabulary size before training neural models**</mark>*, a critical factor in the efficiency and effectiveness of these systems.

{% embed url="<https://arxiv.org/abs/1808.06226>" %}
SentencePiece
{% endembed %}

### <mark style="color:purple;">Understanding SentencePiece</mark>

SentencePiece is a language-independent subword tokenizer and detokenizer, engineered for neural-based text processing.&#x20;

Unlike conventional tokenizers, it doesn't rely on whitespaces for tokenization, making it versatile for languages like Chinese and Japanese.&#x20;

It implements subword units, such as byte-pair-encoding (BPE) and unigram language models, directly from raw sentences. This approach ensures that important words are captured within a fixed vocabulary list, minimising redundancy.

### <mark style="color:purple;">The Role of Tokenization</mark>

Tokenization, the process of breaking down text into words or subwords, is fundamental in NLP.&#x20;

SentencePiece excels in splitting words into subwords, capturing frequent and diverse subwords within a predetermined vocabulary size. &#x20;

#### <mark style="color:green;">**Example Code**</mark>

```python
import sentencepiece as spm

# Initialize SentencePiece
sp = spm.SentencePieceProcessor(model_file='your_model.model')

# Encode text into subwords
encoded_text = sp.encode_as_pieces('This is a sample text.')
print(encoded_text)

# Decode subwords back into text
decoded_text = sp.decode_pieces(encoded_text)
print(decoded_text)
```

#### <mark style="color:green;">Importance of Vocabulary Size Limit</mark>

Setting a vocabulary size limit is vital in preventing the inclusion of rare or complex words that may not be beneficial as separate vectors. This balance is key to efficient and effective models.

### <mark style="color:purple;">Key Components of SentencePiece</mark>

SentencePiece comprises four primary components:

1. <mark style="color:green;">**Normalizer**</mark><mark style="color:green;">:</mark> Standardises words into equivalent NFKC Unicode.
2. <mark style="color:green;">**Trainer**</mark><mark style="color:green;">:</mark> Builds vocabulary based on subword components using BPE and unigram language models.
3. <mark style="color:green;">**Encoder and Decoder**</mark><mark style="color:green;">:</mark> Handle encoding and decoding processes, ensuring lossless tokenization.

### <mark style="color:purple;">Key technical aspects</mark>

#### <mark style="color:green;">Subword segmentation</mark>

SentencePiece implements two subword segmentation algorithms - byte-pair-encoding (BPE) and unigram language model. These algorithms allow the tokenizer to break down words into smaller units (subwords) to reduce the vocabulary size and handle out-of-vocabulary words effectively.

#### <mark style="color:green;">Language independence</mark>

SentencePiece can directly train subword models from raw sentences without relying on language-specific pre-tokenization. This enables the creation of purely end-to-end and language-independent NMT systems.

#### <mark style="color:green;">Lossless tokenization</mark>

SentencePiece treats the input text as a sequence of Unicode characters, including whitespace, which is escaped with a meta symbol (e.g., "\_"). This allows for reversible encoding and decoding without losing information, making the process language-agnostic.

#### <mark style="color:green;">Vocabulary management</mark>

SentencePiece manages the vocabulary-to-id mapping, enabling direct conversion of text into an id sequence and vice versa. This is particularly useful for NMT systems, as their input and output are typically id sequences.

#### <mark style="color:green;">Normalization</mark>

SentencePiece includes a normalizer module that canonicalizes semantically-equivalent Unicode characters, ensuring consistent input for the subword model training.

The paper demonstrates that SentencePiece can achieve comparable accuracy to direct subword training from raw sentences in an English-Japanese NMT task.&#x20;

By providing a simple, language-independent, and reversible tokenization process, SentencePiece aims to standardise and simplify the preprocessing and postprocessing steps in neural network-based NLP systems.

### <mark style="color:purple;">SentencePiece makes the process of tokenization easier in several ways</mark>

#### <mark style="color:green;">**Language Independence**</mark>

SentencePiece is designed to be language-independent, meaning it can be applied to any language without requiring language-specific knowledge or preprocessing. This is particularly useful in multilingual NLP tasks or when working with low-resource languages. By treating the input text as a sequence of Unicode characters and directly learning subword units from raw sentences, SentencePiece eliminates the need for language-specific tokenization rules or tools.

#### <mark style="color:green;">Vocabulary Size Management</mark>

One of the key challenges in tokenization is managing the vocabulary size. A large vocabulary can lead to increased model complexity and computational costs, while a small vocabulary may not capture important words or subwords. SentencePiece allows you to specify a desired vocabulary size, and it automatically learns the most frequent and informative subwords to include in the vocabulary. This helps in striking a balance between model efficiency and expressiveness.

#### <mark style="color:green;">Subword Segmentation</mark>

SentencePiece implements subword segmentation algorithms, such as byte-pair encoding (BPE) and unigram language model, which break down words into smaller units (subwords). This approach has several advantages:

* It reduces the vocabulary size by representing rare or out-of-vocabulary words as combinations of subwords.
* It captures morphological and semantic information within words, as subwords often correspond to meaningful units like prefixes, suffixes, or roots.
* It enables the model to handle unseen words by composing them from learned subwords.

#### <mark style="color:green;">Reversibility and Consistency</mark>

SentencePiece provides a lossless tokenization process, meaning that the original text can be perfectly reconstructed from the tokenized representation. It achieves this by treating whitespace and other special characters as separate tokens and escaping them with a meta symbol. This reversibility ensures that no information is lost during tokenization and detokenization, making it easier to integrate SentencePiece into existing NLP pipelines.

#### <mark style="color:green;">Simplicity and Ease of Use</mark>

SentencePiece offers a simple and intuitive API for tokenization and detokenization. It provides straightforward methods for training subword models, encoding text into subword sequences, and decoding subword sequences back into text. The library is well-documented and comes with pre-trained models for various languages, making it easy to get started with tokenization tasks.

#### <mark style="color:green;">Reproducibility</mark>

SentencePiece promotes reproducibility by providing a standardised and deterministic tokenization process.&#x20;

Given the same input text and trained model, SentencePiece guarantees consistent tokenization results across different platforms and implementations. This is crucial for reproducible research and ensures that models trained using SentencePiece can be easily shared and deployed.


# Tokenization explore

Tokenization is a crucial process in natural language processing (NLP) and large language models (LLMs).

It involves breaking down text into smaller units called tokens, which can be words, subwords, or even characters.&#x20;

The choice of tokenization strategy can significantly impact the performance and behavior of language models.&#x20;

Here's a detailed knowledge document on tokenization and best practices, with relevant code examples.

### <mark style="color:purple;">**Understanding Tokenization**</mark>

Tokenization is the process of converting text into a sequence of tokens that can be processed by language models. Language models operate on numerical representations of text, and tokenization is the first step in this conversion process.&#x20;

The tokens are then *<mark style="color:yellow;">**mapped to unique numerical values (token IDs)**</mark>* that the model can understand.

### <mark style="color:purple;">**Types of Tokenization**</mark>

There are several tokenization strategies, each with its own advantages and trade-offs:&#x20;

#### <mark style="color:green;">**Word-level Tokenization**</mark>

This is the most straightforward approach, where each word is treated as a separate token. However, this can lead to large vocabulary sizes, especially for morphologically rich languages, and it cannot handle out-of-vocabulary (OOV) words or misspellings.

#### <mark style="color:green;">**Character-level Tokenization**</mark>

In this approach, individual characters are treated as tokens. While this allows for handling OOV words and misspellings, it can result in very long sequences, making it computationally expensive for language models.

#### <mark style="color:green;">**Subword Tokenization**</mark>

**T**his is a popular approach that strikes a balance between word-level and character-level tokenization. It breaks down words into smaller units called subwords or wordpieces. This helps reduce vocabulary size and handle OOV words while maintaining context and meaning.

#### <mark style="color:green;">**Byte-Pair Encoding (BPE)**</mark>

BPE is a subword tokenization technique that iteratively merges the most frequent pairs of bytes or characters in the training data to create a vocabulary of subword units.&#x20;

This approach is widely used in state-of-the-art language models like GPT and BERT.

### <mark style="color:purple;">**Best Practices for Tokenization**</mark>

#### <mark style="color:green;">**Use Subword Tokenization**</mark>

For most NLP tasks, subword tokenization techniques like BPE are recommended as they strike a good balance between vocabulary size, handling OOV words, and maintaining context.

#### <mark style="color:green;">Train Tokenizer on Relevant Data</mark>

Train your tokenizer on a corpus that is representative of the data you will be using for your NLP task. This ensures that the tokenizer can handle domain-specific vocabulary and abbreviations effectively.&#x20;

#### <mark style="color:green;">**Preprocess Text**</mark>

Before tokenization, preprocess the text by handling special characters, contractions, and other text normalization steps specific to your task or language.

#### <mark style="color:green;">**Handle Casing**</mark>

Decide whether to preserve or normalize the casing of text before tokenization. Some models are case-sensitive, while others are not

#### <mark style="color:green;">**Control Vocabulary Size**</mark>

When using subword tokenization, you can control the maximum vocabulary size by adjusting the tokenizer's parameters, such as the number of merge operations in BPE or the target vocabulary size

#### <mark style="color:green;">**Tokenize at Inference Time**</mark>

For production scenarios, tokenize the input text at inference time, not during training. This ensures consistency between the tokenization used during training and inference.

### <mark style="color:purple;">**Code Examples**</mark>

<mark style="color:green;">**Word-level Tokenization with NLTK**</mark>

```python
import nltk

text = "This is an example sentence."
tokens = nltk.word_tokenize(text)
print(tokens)
```

Output: `['This', 'is', 'an', 'example', 'sentence', '.']`&#x20;

<mark style="color:green;">**Subword Tokenization with Hugging Face Tokenizers**</mark>

```python
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
text = "This is an example sentence."
tokens = tokenizer.tokenize(text)
print(tokens)
```

Output: `['This', 'is', 'an', 'example', 'sentence', '.']`&#x20;

<mark style="color:green;">**BPE Tokenization with SentencePiece**</mark>

```python
import sentencepiece as spm

# Train a BPE tokenizer
spm.SentencePieceTrainer.Train('--input=data.txt --model_prefix=bpe --vocab_size=10000')

# Load the trained tokenizer
sp = spm.SentencePieceProcessor()
sp.Load('bpe.model')

text = "This is an example sentence."
tokens = sp.EncodeAsPieces(text)
print(tokens)
```

Output: `['▁This', '▁is', '▁an', '▁example', '▁sentence', '.']`

### <mark style="color:purple;">**Tokenization in Large Language Models**</mark>

Large language models like GPT, BERT, and T5 employ advanced tokenization strategies like BPE or WordPiece to handle large vocabularies and OOV words effectively.&#x20;

These models often provide pretrained tokenizers that can be easily loaded and used for tokenization, as shown in the Hugging Face example above.

### <mark style="color:purple;">**Conclusion**</mark>

Tokenization is a critical step in NLP and language modelling, and choosing the right tokenization strategy can significantly impact model performance and behavior.&#x20;

Subword tokenization techniques like BPE are generally recommended as they offer a good balance between vocabulary size, handling OOV words, and maintaining context.&#x20;

Additionally, following best practices like training tokenizers on relevant data, preprocessing text, and controlling vocabulary size can further improve the effectiveness of tokenization for your specific NLP task.


# Tokenizer Choice For LLM Training: Negligible or Crucial?

This  <mark style="color:blue;">**March 2024**</mark> paper investigates the impact of tokenizer choice on the performance of Large Language Models (LLMs), particularly in the context of mono- and multilingual models.&#x20;

The authors argue that while various factors such as dataset composition, model architecture, and pretraining objectives have been extensively studied, the influence of tokenizers remains underexplored.

The paper provides an overview of various tokenization approaches, including word, subword, and character tokenization.&#x20;

It also discusses the usage of tokenizers in encoder and decoder models, highlighting the lack of extensive studies on the extrinsic tokenizer performance in a monolingual and multilingual setting with a focus on decoder-only models.

{% embed url="<https://arxiv.org/abs/2310.08754>" %}
Tokenizer Choice For LLM Training: Negligible or Crucial?
{% endembed %}

### <mark style="color:purple;">Definitions First!</mark>

### <mark style="color:green;">Intrinsic Evaluation</mark>

Intrinsic evaluation assesses the performance of a tokenizer *<mark style="color:yellow;">**based on its inherent properties and the characteristics of its output**</mark>*, without considering its impact on downstream tasks or models. In the context of the study, the intrinsic evaluation focused on three main metrics:

<mark style="color:blue;">**Fertility**</mark>

Fertility is a measure of the average number of tokens required to represent a word or document after tokenization.&#x20;

It is calculated by dividing the total number of tokens in a tokenized dataset by the total number of words in the original dataset.&#x20;

A lower fertility score indicates that the tokenizer is more efficient in representing the text with fewer tokens.&#x20;

High fertility can lead to increased sequence lengths, which can impact the model's ability to learn long-range dependencies and increase computational costs during training and inference.

<mark style="color:blue;">**Parity**</mark>

Parity assesses how fairly a tokenizer treats equivalent sentences across different languages.&#x20;

It measures the consistency of the tokenizer in representing similar semantic content across languages.&#x20;

In the study, parity was calculated using the FLORES-200 parallel corpus, which contains the same sentences human-translated into 200 languages. The closer the parity scores are to 1, the more consistent the tokenizer is in handling different languages.

<mark style="color:blue;">**Vocabulary Overlap**</mark>

Vocabulary overlap measures the similarity between the vocabularies of different tokenizers.&#x20;

It is calculated by comparing the number of shared tokens between two tokenizers' vocabularies.&#x20;

A high vocabulary overlap indicates that the tokenizers generate similar subword units and may have similar performance characteristics.

### <mark style="color:green;">Extrinsic Evaluation</mark>

Extrinsic evaluation *<mark style="color:yellow;">**assesses the performance of a tokenizer based on its impact on downstream tasks or models**</mark>*.

In the study, the extrinsic evaluation focused on two main aspects:

#### <mark style="color:blue;">Downstream Performance</mark>

The authors trained language models using different tokenizers and evaluated their performance on a wide range of downstream tasks, such as natural language inference, question answering, reading comprehension, and text classification.&#x20;

The models' performance on these tasks was used to assess the effectiveness of the tokenizers in enabling the models to learn meaningful representations and solve real-world problems.

#### <mark style="color:blue;">Computational Costs</mark>

The study evaluated the computational costs associated with each tokenizer when used in a specific language model.&#x20;

The computational costs were measured in terms of the average number of floating-point operations (FLOPs) required to process a single word during training.&#x20;

This metric takes into account the tokenizer's fertility and the model's architecture, such as the number of layers, hidden size, and vocabulary size.&#x20;

Lower computational costs indicate that the tokenizer is more efficient in terms of resource utilization during training.

### <mark style="color:green;">Distinction between intrinsic and extrinsic evaluation</mark>

The distinction between intrinsic and extrinsic evaluation is important because a tokenizer's intrinsic performance metrics, such as fertility and parity, may not always directly correlate with its impact on downstream tasks.&#x20;

A tokenizer with good intrinsic performance may not necessarily lead to better downstream performance, and vice versa.&#x20;

Therefore, *<mark style="color:yellow;">**considering both intrinsic and extrinsic evaluation metrics provides a more comprehensive understanding of a tokenizer's effectiveness and its suitability for specific applications**</mark>*.

By conducting both intrinsic and extrinsic evaluations, the study aimed to provide insights into the relationship between tokenizer characteristics and their impact on language model performance, helping researchers and practitioners make informed decisions when selecting or designing tokenizers for their specific use cases.

### <mark style="color:purple;">Key points and findings</mark>

<mark style="color:blue;">**Comprehensive study:**</mark> The authors trained 24 mono- and multilingual LLMs at a 2.6B parameter scale, ablating different tokenizer algorithms and parameterizations to assess their impact on downstream performance.

<mark style="color:blue;">**Significant impact:**</mark> The study reveals that tokenizer choice can significantly affect the model's downstream performance and training costs.

<mark style="color:blue;">**Questionable evaluation metrics:**</mark> Common tokenizer evaluation metrics like fertility and parity are not always predictive of the model's downstream performance, making them unreliable proxies.

<mark style="color:blue;">**Multilingual tokenizers:**</mark> Multilingual tokenizers trained on the five most frequent European languages require vocabulary size increases of factor three compared to English tokenizers.

<mark style="color:blue;">**English-centric tokenizers:**</mark> Applying English-centric tokenizers to multilingual LLMs results in severe downstream performance degradation and additional training costs of up to 68% due to an inefficient tokenization vocabulary.

### <mark style="color:purple;">The approach taken in the study</mark>

#### <mark style="color:green;">Dataset creation</mark>

The authors created two datasets (monolingual English and multilingual) with 70B words each, ensuring that the mixture proportions of data domains were consistent between the tokenizer and model training datasets to avoid domain shift.

#### <mark style="color:green;">Tokenizer training</mark>

They trained 24 different tokenizers using BPE and Unigram algorithms, varying the language composition, vocabulary size, and tokenizer library (Huggingface and SentencePiece).

#### <mark style="color:green;">Model training</mark>

For each of the 24 trained tokenizers, they trained a 2.6B transformer-based decoder-only model on up to 52B tokens, following a specific scaling law. They also trained monolingual and multilingual baseline models using the pre-trained GPT-2 tokenizer.

#### <mark style="color:green;">Evaluation</mark>

The authors conducted both intrinsic and extrinsic evaluations of the tokenizers.&#x20;

The intrinsic evaluation assessed the tokenizers' output based on fertility, parity, and vocabulary overlap, while the extrinsic evaluation measured the impact of tokenizers on the model's downstream performance and computational costs.

### <mark style="color:purple;">To apply these ideas in practice, researchers and practitioners can</mark>

1. Evaluate and compare different tokenizer algorithms, configurations, and libraries for their specific use case, *<mark style="color:yellow;">**considering both intrinsic and extrinsic performance metrics**</mark>*.
2. *<mark style="color:yellow;">**Design and train custom tokenizers tailored to the target languages and domains**</mark>* of their LLMs, ensuring optimal downstream performance and computational efficiency.
3. Use the provided *<mark style="color:yellow;">**computational cost formula to estimate the impact of tokenizer choice**</mark>* and model architecture on training costs, and optimize their setups accordingly.

By applying these insights and approaches, practitioners can develop more efficient and effective LLMs, particularly in multilingual settings, ultimately leading to better performance on downstream tasks and more accessible language technology for a wider range of users.

### <mark style="color:purple;">Key insights that can guide your choice or development of a tokenizer</mark>

#### <mark style="color:green;">Monolingual vs. Multilingual</mark>

If you are working with multilingual data, it is crucial to use a multilingual tokenizer.&#x20;

Using a monolingual tokenizer on multilingual data leads to significantly higher fertility and parity scores, indicating inefficient tokenization.&#x20;

Multilingual tokenizers perform only slightly worse on English documents compared to monolingual English tokenizers, making them a better choice for multilingual settings.

#### <mark style="color:green;">Vocabulary Size</mark>

The optimal vocabulary size depends on the language(s) you are working with.&#x20;

For monolingual English settings, smaller vocabulary sizes (around 33k-50k) tend to perform better.&#x20;

However, for multilingual settings, larger vocabulary sizes (up to 100k) generally yield better downstream performance, especially for non-Germanic languages. Consider the trade-off between performance and computational costs when choosing the vocabulary size.

#### <mark style="color:green;">Tokenizer Library</mark>

The choice of tokenizer library (e.g., Huggingface vs. SentencePiece) can impact the downstream performance.&#x20;

In the study, *<mark style="color:yellow;">**SentencePiece's BPE implementation generally outperformed Huggingface's BPE**</mark>* in both monolingual and multilingual settings.&#x20;

Consider the differences in pre- and post-processing steps between libraries when selecting a tokenizer.

#### <mark style="color:green;">Tokenizer Algorithm</mark>

The choice between BPE and Unigram algorithms may depend on the target language(s).&#x20;

In the study, Germanic languages (German and English) benefited more from BPE, while Romanic languages (French and Spanish) benefited more from Unigram.&#x20;

Consider the linguistic properties of your target language(s) when choosing the tokenizer algorithm.

#### <mark style="color:green;">Computational Costs</mark>

Larger vocabulary sizes generally increase computational costs, even if they lead to lower fertility scores.&#x20;

When choosing a tokenizer, consider the trade-off between downstream performance and computational costs, especially if you need to process a fixed set of documents during training.

#### <mark style="color:green;">Pre-trained Tokenizers</mark>

Using pre-trained tokenizers (e.g., GPT-2) for multilingual models may lead to suboptimal performance. It is better to train a custom tokenizer tailored to your specific language(s) and domain.

### <mark style="color:purple;">Conclusion</mark>

The study provides valuable insights into the impact of tokenizer choice on the downstream performance of language models, particularly in monolingual and multilingual settings.&#x20;

The findings highlight the importance of training tokenizers with a balanced share across languages to achieve comparable low fertility and parity scores, which has significant implications for computational costs and the model's ability to learn long-range dependencies.

The study demonstrates that the tokenizer choice can significantly impact the model's downstream performance, with the BPE algorithm performing well in both mono- and multilingual settings.&#x20;

For English, a vocabulary size of 33k is sufficient, while multilingual models based on the five considered languages require up to three times larger vocabulary sizes. Additionally, the SentencePiece library outperforms the Huggingface tokenizer library.

Interestingly, the study finds no clear correlation between intrinsic and extrinsic tokenizer performance, suggesting that the correlation is rather task-specific. A small fertility value might be a necessary condition for good downstream performance but not a sufficient one.


# Getting the most out of your tokenizer for pre-training and domain adaptation

Tokenization is often an understudied and neglected component in the development of models.

In this <mark style="color:blue;">**February 2024**</mark> paper, the authors highlight that most published works use a single tokenizer for all experiments, often borrowed from another model, without performing rigorous analysis or ablations to optimise the tokenization process.&#x20;

Furthermore, when fine-tuning a pre-trained LLM for a specific task or domain, the tokenizer is generally kept unchanged, leading to sub-optimal performance and efficiency.

The authors argue that the <mark style="color:yellow;">**size of the tokenizer's vocabulary**</mark>, the <mark style="color:yellow;">**pre-tokenization regular expression**</mark>, and the <mark style="color:yellow;">**training data**</mark> used for the tokenizer can significantly impact the model's generation speed, effective context size, memory usage, and downstream performance.

To address this issue, the authors train specialised <mark style="color:blue;">**Byte-Pair Encoding (BPE)**</mark> code tokenizers and conduct extensive ablations to study the impact of tokenizer design on the performance of LLMs for code generation tasks such as HumanEval and MBPP.&#x20;

They provide recommendations for selecting appropriate tokenizer hyper-parameters and suggest switching the tokenizer when fine-tuning a pre-trained LLM.

The experiments are performed on models trained from scratch and on pre-trained models, verifying the applicability of their findings to a wide range of use-cases.&#x20;

The authors find that when fine-tuning on more than 50 billion tokens, it is possible to specialise the tokenizer of a pre-trained LLM to obtain significant gains in generation speed and effective context size.

{% embed url="<https://arxiv.org/abs/2402.01035>" %}
Getting the most out of your tokenizer for pre-training and domain adaptation
{% endembed %}

<figure><img src="/files/QnUIE0pay3LhD2ii5jBJ" alt=""><figcaption><p>Three ways to increase in-domain compression in a BPE tokenizer with their respective trade-offs</p></figcaption></figure>

### <mark style="color:purple;">Aspects of tokenizer design</mark>

### <mark style="color:green;">Compression Trade-offs</mark>

<mark style="color:blue;">**Training Data:**</mark> Using data sampled from the target domain (e.g., code) will increase compression for that domain.

<mark style="color:blue;">**Pre-tokenization Scheme:**</mark> The regular expression used to split the text before applying BPE affects compression. Splitting on whitespaces prevents BPE from merging across words, leading to shorter tokens and worse compression.

<mark style="color:blue;">**Vocabulary Size:**</mark> A larger vocabulary size leads to higher compression but increases computational and memory costs.

### <mark style="color:green;">Compression Metrics</mark>

<mark style="color:blue;">**Normalized Sequence Length (NSL):**</mark> Measures the average tokenized sequence length of a tokenizer compared to a baseline (Llama tokenizer). An NSL of 0.75 means the tokenizer uses 25% fewer tokens on average.&#x20;

<mark style="color:blue;">**Bytes per Token:**</mark> Calculated by dividing the number of UTF-8 bytes by the number of tokens, providing another measure of compression.

### <mark style="color:green;">BPE Algorithm and Implementation</mark>

The authors use the BPE tokenization algorithm, implemented with the HuggingFace tokenizers library, which supports regular expression-based pre-tokenization and handles special formatting characters better.

### <mark style="color:green;">Impact of Training Data</mark>

Unsurprisingly, training the tokenizer on data from the target domain (code, English, multilingual) improves compression for that domain. Training on a mix of all three leads to the best overall compression.

### <mark style="color:green;">Vocabulary Size</mark>

<mark style="color:blue;">Compression vs. Vocabulary Size:</mark> Larger vocabularies improve compression, but gains diminish exponentially as the vocabulary size increases.&#x20;

<mark style="color:blue;">Inference Optimal Vocabulary Size</mark>: The authors calculate the optimal vocabulary size for inference time by considering the trade-off between compression gains and additional computation costs.&#x20;

<mark style="color:blue;">Memory Optimal Vocabulary Size:</mark> The authors derive an equation to find the memory-optimal vocabulary size, considering the model size, sequence length, batch size, and the memory savings from reduced attention cache size due to compression.

### <mark style="color:purple;">Ideas for creating tokenizers in different fields</mark>

<mark style="color:green;">**Biomedical Domain**</mark>

* Develop a tokenizer tailored for biomedical literature, such as research papers, clinical notes, and scientific reports.
* Train the tokenizer on a large corpus of biomedical texts, incorporating domain-specific vocabulary, abbreviations, and naming conventions.
* Use pre-tokenization schemes that preserve important biomedical entities, such as gene names, protein structures, and chemical compounds.
* Integrate domain-specific knowledge bases or ontologies to improve tokenization accuracy and semantic understanding.

<mark style="color:green;">**Legal and Regulatory Domain**</mark>

* Create a tokenizer specifically designed for legal documents, contracts, regulations, and legislative texts.
* Train the tokenizer on a diverse corpus of legal texts, including case laws, statutes, and regulatory guidelines.
* Implement pre-tokenization rules that preserve legal terminology, citations, and references to specific clauses or sections.
* Incorporate legal ontologies and dictionaries to accurately tokenize complex legal phrases and terms.

<mark style="color:green;">**Financial and Accounting Domain**</mark>

* Develop a tokenizer tailored for financial reports, accounting statements, and market data.
* Train the tokenizer on a corpus of financial documents, including annual reports, balance sheets, and market analyses.
* Implement pre-tokenization rules to handle financial notation, abbreviations, and numerical representations.
* Integrate domain-specific knowledge bases or lexicons to accurately tokenize financial terms and concepts.

<mark style="color:green;">**Cybersecurity Domain**</mark>

* Create a tokenizer specifically designed for cybersecurity logs, incident reports, and technical documentation.
* Train the tokenizer on a corpus of cybersecurity-related texts, including system logs, vulnerability reports, and security guidelines.
* Implement pre-tokenization rules to preserve technical terminology, IP addresses, and other relevant cybersecurity entities.
* Integrate domain-specific knowledge bases or ontologies to accurately tokenize cybersecurity-related terms and concepts.

<mark style="color:green;">**Social Media and Conversational Data**</mark>

* Develop a tokenizer tailored for social media data, such as tweets, online forums, and chat conversations.
* Train the tokenizer on a diverse corpus of social media data, including slang, abbreviations, and internet-specific language.
* Implement pre-tokenization rules to handle emoticons, hashtags, and other social media-specific constructs.
* Incorporate language models or lexicons specific to social media and conversational data to improve tokenization accuracy.

### <mark style="color:purple;">Key Takeaways</mark>

Based on the experiments described in this section, here are the key conclusions and best practices regarding tokenizer construction and usage:

#### <mark style="color:green;">Tokenizer Impact on Performance</mark>

* Changing the tokenizer of a pre-trained LLM during fine-tuning can have a negligible impact on downstream performance, provided that the fine-tuning is done on a sufficiently large amount of data (50 billion tokens or more).
* The authors demonstrate that models fine-tuned with alternative tokenizers like GPT-4 and Punct can achieve competitive or even better performance compared to models using the original Llama tokenizer.

#### <mark style="color:green;">Vocabulary Size and Performance</mark>

* The experiments suggest that the vocabulary size of the tokenizer (within the tested range of 32k to 256k) has a minimal impact on the downstream performance of the LLM.
* The authors found no statistically significant correlation between vocabulary size and performance metrics like Pass\@1 and Pass\@100 on code generation tasks.

#### <mark style="color:green;">Tokenizer Update Methods</mark>

* Using techniques like Fast Vocabulary Transfer (FVT) to initialize the new tokenizer's embeddings from the pre-trained model leads to noticeable performance improvements compared to not using FVT.
* Extending an existing tokenizer (e.g., Llama) by adding domain-specific tokens provides only small gains compared to using a completely different tokenizer like GPT-4.

#### <mark style="color:green;">Tokenizer Choice and Compression</mark>

* While highly compressed tokenizers like Identity can offer significant compression benefits, they may result in deteriorated downstream performance on code generation tasks.
* Tokenizers like Punct and GPT-4, which strike a balance between compression and preserving syntactic and semantic information, can achieve both better performance and better compression compared to the Llama tokenizer.

#### <mark style="color:green;">Scaling to Larger LLMs</mark>

* The authors demonstrate that their findings regarding tokenizer switching and its negligible impact on performance hold true for larger LLMs like Llama 2 7B when fine-tuned on a sufficient amount of data.

### <mark style="color:purple;">Summary</mark>

In this study, the authors investigated the impact of tokenizer design choices on the performance, compression, and efficiency of large language models (LLMs), with a focus on code generation tasks.&#x20;

Their findings highlight the importance of carefully considering tokenization strategies, as they can significantly influence model capabilities and resource utilization.

Through extensive experimentation, they demonstrated that changing the tokenizer of a pre-trained LLM during fine-tuning can have a negligible impact on downstream performance, provided that the fine-tuning is conducted on a sufficiently large dataset (50 billion tokens or more).&#x20;

This insight opens up opportunities for optimizing tokenizers for specific domains or tasks without sacrificing model accuracy.

Furthermore, their results suggest that the vocabulary size of the tokenizer, within a reasonable range (32k to 256k), has a minimal effect on the LLM's downstream performance. This finding allows for flexibility in balancing compression and memory/compute trade-offs based on the specific requirements of the application.

Overall, this study underscores the importance of carefully considering tokenization strategies in the development and fine-tuning of LLMs, as well as the potential for optimizing tokenizers to enhance model efficiency and domain-specific performance.

### <mark style="color:purple;">References</mark>

1. **01.AI Yi series models (2023)**: Discusses the Yi series of large language models available on the Hugging Face Model Repository, highlighting their advanced capabilities in language understanding.
2. **Ahmad et al. (2021)**: Explores unified pre-training methods for program understanding and generation, demonstrating a method to enhance model efficiency in understanding and generating programmatic content.
3. **Ainslie et al. (2023)**: Describes training generalized multi-query transformer models from multi-head checkpoints, advancing the flexibility of transformer architectures in handling diverse queries simultaneously.
4. **Allal et al. (2023)**: Introduces 'Santacoder', a concept for a playful, yet robust approach to coding assistance, enhancing code-related tasks without aiming for overly ambitious, unattainable goals.
5. **Almazrouei et al. (2023)**: Discusses the Falcon series of open language models, emphasizing the development of accessible and robust language model frameworks.
6. **Anthropic (2023)**: Details the release of 'Claude' by Anthropic, focusing on a new language model that aims to improve ethical considerations and robustness in AI.
7. **Austin et al. (2021)**: Investigates the use of large language models for program synthesis, showing their potential in automating coding tasks and generating programmatic content from high-level descriptions.
8. **Biderman et al. (2023)**: Presents 'Pythia', a suite designed for analyzing large language models across different stages of their training and scaling, aiming to understand their behavior and improve their design.
9. **Black et al. (2022)**: Discusses GPT-NeoX-20B, an open-source autoregressive language model, contributing to the open research and development of scalable language models.
10. **Chen et al. (2021)**: Evaluates large language models trained on code, providing insights into their effectiveness and areas for improvement in programming language understanding.
11. **Chirkova & Troshin (2023)**: Explores subtokenization options for pretraining large language models on source code, aiming to optimize model performance on coding tasks.
12. **Deci (2023)**: Introduces 'Decicoder', a model touted as a new standard in efficient and accurate code generation, enhancing the capabilities of AI in software development.
13. **DeepSeek AI (2023)**: Details a series of code language models, enhancing the tools available for developers and programmers in automated code generation.
14. **Devlin et al. (2019)**: Describes the training of BERT, a foundational model that uses deep bidirectional transformers for improved language understanding, setting a new standard in NLP.
15. **Elsen et al. (2023)**: Announces the release of Persimmon-8B, a language model designed to follow short-form instructions effectively, demonstrating its practical applications.
16. **Forsythe (2023)**: Introduces 'Tokenmonster', a tokenizer and vocabulary trainer designed to improve the efficiency of language processing in Python, Go, and JavaScript.
17. **Fried et al. (2023)**: Presents 'Incoder', a generative model for code infilling and synthesis, aimed at enhancing automated coding tasks by filling in gaps and generating syntactically correct code snippets.
18. **Gee et al. (2022, 2023)**: Discusses methods for fast vocabulary transfer and multi-word tokenization for language model compression and sequence compression, respectively, aiming to enhance model efficiency and manage larger vocabularies effectively.
19. **Gowda & May (2020)**: Investigates the optimal vocabulary size for neural machine translation, providing insights that help improve translation accuracy and efficiency.
20. **Goyal et al. (2023)**: Describes the training of language models with pause tokens, introducing a method to enhance natural language generation by incorporating thoughtful pauses in speech or text.
21. **guidance-ai (2023)**: Details a guidance language designed for controlling large language models, emphasizing the development of more responsive and controllable AI systems.
22. **Jiang et al. (2023)**: Discusses 'Mistral 7b', a model designed for multi-turn program synthesis, enhancing the interactive capabilities of AI in coding tasks.
23. **Kocetkov et al. (2022)**: Presents 'The Stack', a large dataset of permissively licensed source code, aimed at facilitating the training and development of code-oriented AI models.
24. **Kudo (2018)**: Investigates subword regularization techniques, aiming to improve translation models by managing subword variability effectively.
25. **Kudo & Richardson (2018)**: Introduces 'Sentencepiece', a tokenizer that simplifies text processing by providing a consistent subword tokenization method


# TokenMonster

TokenMonster is an ungreedy subword tokenizer and vocabulary generator designed to improve the efficiency and performance of language models.&#x20;

It selects an optimal vocabulary for a given dataset, resulting in up to 37.5% fewer tokens required to represent text compared to other modern tokenizing methods. This allows for faster inference, training, and longer text generation.

{% embed url="<https://github.com/alasdairforsythe/tokenmonster>" %}

### <mark style="color:purple;">Key technical features include</mark>

* Ungreedy tokenization algorithm that follows up to 6 parallel branches
* Supports 5 optimisation modes: unfiltered, clean, balanced, consistent, strict
* Uses capcode marker tokens to encode uppercasing and forward delete
* Identifies words, subwords, common phrases, and figures of speech
* Achieves up to 7 characters per token depending on vocabulary size and optimization mode
* Provides 422 pretrained vocabularies and tools to train custom vocabularies
* Implementations available in Go, Python, and JavaScript

### <mark style="color:purple;">Using TokenMonster in Practice</mark>

1. Choose a suitable pretrained vocabulary based on your dataset (e.g., code, English, fiction), desired vocabulary size, and optimization mode. Alternatively, train a custom vocabulary using the provided tools.
2. Install the TokenMonster library in your preferred language (Go, Python, or JavaScript).
3. Load the selected vocabulary:

```python
import tokenmonster
vocab = tokenmonster.load("englishcode-32000-consistent-v1")
```

Tokenize your text using the loaded vocabulary:

```python
tokens = vocab.tokenize("This is a test.")
```

Integrate the tokenized text into your language model training or inference pipeline to benefit from the optimized vocabulary and improved efficiency.

By using TokenMonster, you can potentially reduce the vocabulary size of your language model by 50-75% while maintaining or improving performance.&#x20;

This frees up resources that can be used to make the model smarter and faster.&#x20;

The ungreedy tokenization algorithm and carefully selected vocabularies enable more efficient usage of embeddings and simpler grammar for the model to learn.

An explanation

<details>

<summary><mark style="color:green;"><code>tokenmaster.py</code> codebase</mark></summary>

The `tokenmaster.py` codebase is a Python library for the TokenMonster tokenizer.

It provides an interface to load, modify, and use TokenMonster vocabularies for efficient tokenization and detokenization of text.

**Key components and usage**

1. Loading a vocabulary:
   * Use `tokenmonster.load(path)` to load a vocabulary from a file, URL, or pre-built vocabulary name.
   * For multiprocessing, use `tokenmonster.load_multiprocess_safe(path)` to load the vocabulary safely.
2. Tokenizing text:
   * Use `vocab.tokenize(text)` to tokenize a string or a list of strings into token IDs.
   * The method returns a numpy array or a list of numpy arrays containing the token IDs.
3. Decoding tokens:
   * Use `vocab.decode(tokens)` to decode a single token ID or a list of token IDs back into a string.
   * For decoding token streams sequentially, create a decoder object using `decoder = vocab.decoder()` and use `decoder.decode(tokens)` to decode tokens incrementally.
4. Modifying the vocabulary:
   * Use `vocab.modify()` to add or delete tokens, resize the vocabulary, enable/disable the UNK token, or reset token IDs.
   * Modifications can also be made using individual methods like `add_token()`, `delete_token()`, `add_special_token()`, etc.
   * After modifying the vocabulary, save it using `vocab.save(filename)`.
5. Accessing vocabulary information:
   * Use `vocab.get_dictionary()` to retrieve a dictionary of all tokens in the vocabulary.
   * Access properties like `vocab.vocab_size`, `vocab.unk_token_id()`, `vocab.capcode()`, `vocab.charset()`, etc., to get information about the vocabulary.

<mark style="color:green;">**To build your own tokenizer**</mark>

1. Prepare a dataset of text that represents the domain you want to tokenize.
2. Use the `tokenmonster` library to train a new vocabulary on your dataset:
   * Create a new vocabulary using `vocab = tokenmonster.new(yaml)`, where `yaml` is a YAML string defining the vocabulary configuration.
   * Customise the vocabulary configuration, specifying the desired vocabulary size, optimization mode, and other parameters.
   * Save the trained vocabulary using `vocab.save(filename)`.
3. Use the trained vocabulary to tokenize and detokenize text in your application:
   * Load the saved vocabulary using `vocab = tokenmonster.load(filename)`.
   * Tokenize text using `tokens = vocab.tokenize(text)`.
   * Decode tokens back into text using `decoded_text = vocab.decode(tokens)`.

The key inputs for building your own tokenizer are:

* A representative dataset of text for training the vocabulary.
* A YAML configuration file specifying the vocabulary parameters (size, optimization mode, etc.).

By following these steps and leveraging the `tokenmonster` library, you can build a custom tokenizer optimized for your specific domain and use case.

Remember to handle any errors and exceptions appropriately, and refer to the documentation and examples provided in the TokenMonster repository for more detailed guidance on using the library.

</details>

### <mark style="color:purple;">Case Study: Training a TokenMonster Tokenizer for Medical Text</mark>

Objective: Build a domain-specific tokenizer for medical text to improve the efficiency and accuracy of a fine-tuned language model for medical question answering.

### <mark style="color:blue;">Process</mark>

#### <mark style="color:green;">Data Collection</mark>

* Gather a large corpus of medical text, such as medical research papers, clinical notes, and medical textbooks.
* Ensure the dataset is representative of the medical domain and covers various medical specialties and terminology.

#### <mark style="color:green;">Data Preprocessing</mark>

* Clean the dataset by removing any irrelevant information, such as headers, footers, or metadata.
* Normalize the text by handling special characters, converting to lowercase, and addressing any domain-specific formatting.

#### <mark style="color:green;">Vocabulary Training</mark>

* Prepare a YAML configuration file specifying the desired vocabulary size (e.g., 32,000 tokens), optimization mode (e.g., "consistent"), and any additional settings.
* Create a new TokenMonster vocabulary using the preprocessed medical text dataset:

```python
import tokenmonster
yaml_config = """
vocab_size: 32000
optimization_mode: consistent
"""
vocab = tokenmonster.new(yaml_config)
```

#### <mark style="color:green;">Train the vocabulary on the medical text dataset</mark>

```python
medical_text = load_medical_dataset()
vocab.tokenize(medical_text)
```

#### <mark style="color:green;">Save the trained vocabulary</mark>

```python
vocab.save("medical_tokenizer.vocab")
```

### <mark style="color:blue;">Fine-tuning the Language Model</mark>

#### <mark style="color:green;">Load the trained TokenMonster vocabulary</mark>

```python
medical_tokenizer = tokenmonster.load("medical_tokenizer.vocab")
```

#### <mark style="color:green;">Tokenize the medical text dataset using the trained tokenizer:</mark>

```python
tokenized_medical_text = medical_tokenizer.tokenize(medical_text)
```

* Fine-tune a pre-trained language model (e.g., BERT, RoBERTa) on the tokenized medical text dataset.
* During fine-tuning, use the TokenMonster vocabulary to tokenize the input text and convert the model's output tokens back to text.

### <mark style="color:blue;">Inference and Evaluation</mark>

* Load the fine-tuned medical language model.
* For inference, tokenize the input medical questions using the TokenMonster tokenizer:

```python
equestion = "What are the symptoms of pneumonia?"
tokenized_question = medical_tokenizer.tokenize(question)
```

* Feed the tokenized question to the fine-tuned model and obtain the predicted answer tokens.
* Decode the predicted answer tokens back to text using the TokenMonster tokenizer:

```python
predicted_answer_tokens = model.predict(tokenized_question)
predicted_answer = medical_tokenizer.decode(predicted_answer_tokens)
```

* Evaluate the model's performance using appropriate metrics for medical question answering, such as accuracy, F1 score, or BLEU score.

By incorporating TokenMonster into your workflow, you can create domain-specific tokenizers that capture the unique vocabulary and patterns of your target domain. This can lead to improved efficiency and accuracy in fine-tuning language models for specialized tasks, such as medical question answering, legal document analysis, or scientific text generation.

Remember to evaluate the performance of your fine-tuned model and iterate on the tokenizer training and model fine-tuning process to achieve the best results for your specific use case.


# Parameter Efficient Fine Tuning


# P-Tuning

The highly cited "GPT Understands Too" paper first submitted March 2021, introducing P-Tuning

This <mark style="color:blue;">**March 2021**</mark> paper introduced a method called <mark style="color:blue;">**P-Tuning**</mark>.

P-Tuning is aims to improve and stabilise the performance of prompting in natural language tasks by *<mark style="color:yellow;">**using continuous prompt embeddings instead of discrete prompt tokens.**</mark>*&#x20;

The main idea is to *<mark style="color:yellow;">**concatenate learnable continuous prompt embeddings with the input tokens**</mark>* and optimise them through backpropagation to achieve better task performance and reduce the instability caused by discrete prompts.

To add the continuous prompt tokens to the model, you *<mark style="color:yellow;">**modify the embedding layer of the Transformer to include the additional learnable embeddings**</mark>*.  These embeddings are then concatenated with the input token embeddings before being passed through the self-attention layers.

{% embed url="<https://arxiv.org/abs/2103.10385>" %}
"P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks" by Xiao Liu et al
{% endembed %}

### <mark style="color:purple;">How does P-tuning differ to traditional prompting?</mark>

In traditional prompting, you would use *<mark style="color:yellow;">**fixed, manually-created prompts**</mark>* to guide the language model to perform a specific task.&#x20;

For example, if you want the model to answer a question about a country's capital, you might use a prompt like:

"The capital of \[country] is \[answer]."&#x20;

Here, "\[country]" and "\[answer]" are placeholders that will be replaced with the actual country and the model's predicted answer, respectively.

However, *<mark style="color:yellow;">**creating these prompts manually can be time-consuming and may not always lead to the best performance on the task.**</mark>*  This is where P-Tuning comes in.

Instead of using fixed, discrete prompts, P-Tuning introduces learnable, continuous prompt embeddings. These *<mark style="color:yellow;">**embeddings are like a set of "virtual" words that are learned during the training process.**</mark>*&#x20;

They are called "continuous" because they are represented as real-valued vectors, as opposed to discrete tokens like words.

### <mark style="color:purple;">**Here's a simplified step-by-step explanation of how P-Tuning works**</mark>

1. You <mark style="color:yellow;">define a prompt template</mark> that includes placeholders for the input (e.g., the question), the output (e.g., the answer), and the continuous prompt embeddings. These embeddings are randomly initialised at the beginning.
2. The continuous <mark style="color:yellow;">prompt embeddings are added to the embedding layer</mark> of the Transformer model, along with the embeddings of the actual input tokens and output labels.
3. An <mark style="color:yellow;">additional mapping function is used</mark> to map the continuous prompt embeddings to the hidden states of the model. This function can be a simple neural network like an <mark style="color:blue;">**Long Short-Term Memory (LSTM)**</mark> or <mark style="color:blue;">**Multilayer Perceptron (MLP)**</mark>
4. During training, the <mark style="color:yellow;">continuous prompt embeddings are updated based on the model's performance on the task</mark>. The model learns to adjust these embeddings to minimise the task-specific loss, just like it learns to adjust its other parameters.
5. At inference time, the learned <mark style="color:yellow;">continuous prompt embeddings are combined with the input tokens</mark> and fed into the Transformer model to generate predictions.

The key idea is that *<mark style="color:yellow;">**by learning these continuous prompt embeddings, the model can automatically discover the best "prompts" for the task during training**</mark>*, rather than relying on manually-created, fixed prompts.&#x20;

This can lead to better performance and more flexibility in adapting the model to different tasks.

<details>

<summary><mark style="color:blue;"><strong>Definition:</strong></mark> <mark style="color:blue;"><strong>LSTM (Long Short-Term Memory) and MLP (Multilayer Perceptron)</strong></mark> </summary>

LSTM (Long Short-Term Memory) and MLP (Multilayer Perceptron) are two types of neural network architectures that can be used as the mapping function in P-Tuning to transform the continuous prompt embeddings into the hidden states of the model.

<mark style="color:green;">**LSTM (Long Short-Term Memory)**</mark>

* LSTM is a type of recurrent neural network (RNN) architecture designed to handle sequential data and capture long-term dependencies.
* It consists of a unique cell state and multiple gating mechanisms (input gate, forget gate, and output gate) that regulate the flow of information in and out of the cell.
* The cell state acts as a memory unit, allowing the LSTM to selectively remember or forget information over long sequences.
* LSTMs are particularly effective in tasks involving sequential data, such as natural language processing, speech recognition, and time series analysis.
* In the context of P-Tuning, an LSTM can be used to process the continuous prompt embeddings and generate hidden states that capture the contextual information and long-term dependencies within the prompts.

<mark style="color:green;">**MLP (Multilayer Perceptron)**</mark>

* MLP is a feedforward neural network architecture consisting of multiple layers of interconnected nodes (neurons).
* It has an input layer, one or more hidden layers, and an output layer.
* Each neuron in an MLP applies a nonlinear activation function to a weighted sum of its inputs, allowing the network to learn complex nonlinear mappings between the input and output.
* MLPs are versatile and can be used for a wide range of tasks, including classification, regression, and feature learning.
* In the context of P-Tuning, an MLP can be used to transform the continuous prompt embeddings into hidden states by applying a series of linear transformations and nonlinear activations.

Both LSTM and MLP can be used as the mapping function in P-Tuning, depending on the specific requirements of the task and the nature of the prompt embeddings.&#x20;

LSTMs are particularly suitable when the prompts have a sequential structure and capturing long-term dependencies is important.&#x20;

MLPs, on the other hand, are simpler and more straightforward, making them a good choice when the prompts do not have a strong sequential nature or when computational efficiency is a priority.

The choice between LSTM and MLP as the mapping function in P-Tuning ultimately depends on the characteristics of the task, the complexity of the prompts, and the available computational resources.&#x20;

Experimenting with both architectures and comparing their performance can help determine the most suitable choice for a given application.

</details>

### <mark style="color:purple;">What does "concatenate learnable continuous prompt embeddings" mean?</mark>

It means we are combining <mark style="color:yellow;">**two types**</mark> of embeddings:

#### <mark style="color:green;">**Input token embeddings**</mark>

These are the embeddings of the <mark style="color:yellow;">**actual input tokens**</mark> (words or subwords) that represent the text data we want to process.  In the Transformer architecture, each input token is mapped to a dense vector representation (embedding) that captures its semantic meaning.

#### <mark style="color:green;">Learnable continuous prompt embeddings</mark>

These are <mark style="color:yellow;">**additional embeddings**</mark> that are not associated with any specific input token but are learned during the training process.&#x20;

They are called "continuous" because they are *<mark style="color:yellow;">**represented as dense vectors in a continuous space**</mark>*, as opposed to discrete tokens. These embeddings serve as a "prompt" that guides the model to perform better on the specific task.

The process of concatenation involves *<mark style="color:yellow;">**joining these two types of embeddings together to form a single input sequence**</mark>*.&#x20;

The key difference between using learnable continuous prompt embeddings and discrete prompts is that the *<mark style="color:yellow;">**continuous embeddings are optimised through backpropagation during training.**</mark>*&#x20;

This means that the model can learn to adjust these embeddings based on the specific task and the training data, allowing for more flexibility and adaptability. In contrast, discrete prompts are fixed and cannot be optimised during training.

By optimising the continuous prompt embeddings through backpropagation, the model can learn to generate more informative and stable prompts, which can lead to better task performance and reduce the instability caused by manually-crafted discrete prompts.

### <mark style="color:purple;">An example of concatenation in P-Tuning</mark>

Let's break it down into a simple, everyday example to better understand the concept of concatenating input token embeddings and learnable continuous prompt embeddings.

Imagine you're planning a trip and have a list of essential items you need to pack:

* Toothbrush
* Toothpaste
* Shampoo
* Conditioner
* Clothes

These items are like the *<mark style="color:yellow;">**input token embeddings**</mark>* - they are the basic elements you need for your trip.

Now, to make your trip more organised and enjoyable, you decide to add some additional items to your list:

* Travel-sized toothbrush case
* Travel-sized toothpaste tube
* Travel-sized shampoo bottle
* Travel-sized conditioner bottle
* Laundry bag for dirty clothes

These additional items are like the *<mark style="color:yellow;">**learnable continuous prompt embeddings**</mark>* - they enhance and support the basic elements of your trip.

These embeddings *<mark style="color:yellow;">**are not tied to specific words but are learned during the training process to guide the language model in generating relevant and coherent packing list**</mark><mark style="color:yellow;">s.</mark>*

The process of concatenation is like combining these two lists into a single, comprehensive packing list:

\[Travel-sized toothbrush case, Toothbrush, Travel-sized toothpaste tube, Toothpaste, Travel-sized shampoo bottle, Shampoo, Travel-sized conditioner bottle, Conditioner, Laundry bag for dirty clothes, Clothes]

By concatenating the additional items with the essential items, you create a single, organised list that helps you better prepare for your trip.

### <mark style="color:purple;">The process</mark>

The concatenated input sequence, containing both the input token embeddings and the learnable continuous prompt embeddings, is then fed into the language model.

The language model processes this unified input sequence and learns to generate coherent and relevant travel packing lists based on the provided context and prompts.

By concatenating the learnable continuous prompt embeddings with the input token embeddings, P-Tuning allows the language model to leverage both the semantic information from the actual input tokens and the guiding information from the learned prompts.&#x20;

This concatenation helps the model generate more accurate and context-aware outputs for the specific task at hand.

### <mark style="color:purple;">Diagram of Concept from the Paper</mark>

<figure><img src="/files/BlmGMuSy3v4XT5yt0dg8" alt=""><figcaption><p>An example of prompt search for “The capital of Britain is [MASK]”. Given the context (blue zone, “Britain”) and target (red zone, “[MASK]”), the orange zone refer to the prompt. In (a), the prompt generator only receives discrete rewards; on the contrary, in (b) the continuous prompt embeddings and prompt encoder can be optimized in a differentiable way.</p></figcaption></figure>

### <mark style="color:purple;">The key advantages of P-Tuning include</mark>

<mark style="color:blue;">**Improved performance:**</mark> By learning optimal prompt embeddings during training, P-Tuning enables language models to achieve better results on a wide range of natural language understanding tasks.

<mark style="color:blue;">**Increased flexibility:**</mark> P-Tuning allows language models to adapt more effectively to different tasks and domains by learning task-specific prompts, reducing the need for extensive fine-tuning or manual prompt engineering.

<mark style="color:blue;">**Enhanced interpretability:**</mark> The learned continuous prompt embeddings provide insights into the language model's behaviour and the important aspects of the task, making the model's decisions more interpretable and explainable.

<mark style="color:blue;">**Efficient adaptation:**</mark> P-Tuning offers a more efficient way to adapt language models to new tasks, as it focuses on learning prompts rather than modifying the entire model architecture or weights.


# The Power of Scale for Parameter-Efficient Prompt Tuning

This highly cited <mark style="color:blue;">**September 2021**</mark> paper introduced "prompt tuning," a method for adapting large pre-trained language models to perform specific downstream tasks by learning "soft prompts" that condition the model's behaviour.&#x20;

Prompt tuning was one of the first concepts around parameter efficient fine tuning

{% embed url="<https://arxiv.org/abs/2104.08691>" %}
The Power of Scale for Parameter-Efficient Prompt Tuning
{% endembed %}

Instead of using discrete text prompts like GPT-3, the authors propose learning continuous "soft prompts" through backpropagation.  These soft prompts can incorporate signals from labeled examples and outperform GPT-3's few-shot learning.

Through experiments with the T5 model, the authors show that prompt tuning becomes more competitive with model tuning (where all model weights are tuned) as the model size increases.  With billion-parameter models, *<mark style="color:yellow;">**prompt tuning can match the performance of model tuning.**</mark>*

Prompt tuning is *<mark style="color:yellow;">**more parameter-efficient than model tuning**</mark>*, as a single frozen model can be reused for multiple downstream tasks by learning task-specific prompts. This is especially beneficial for large models that are costly to share and serve.

The authors compare prompt tuning to similar approaches like "prefix tuning" (Li and Liang, 2021) and show that prompt tuning alone, without intermediate-layer prefixes or task-specific output layers, is sufficient to be competitive with model tuning.

Prompt tuning has additional benefits, such as better resilience to domain shifts compared to model tuning, and the ability to perform efficient <mark style="color:blue;">**"prompt ensembling"**</mark> by learning multiple prompts for the same task.

### <mark style="color:purple;">Key Features of Prompt Tuning</mark>

#### <mark style="color:green;">Parameter efficiency</mark>

Prompt tuning is highly parameter-efficient compared to other methods. It requires less than 0.01% task-specific parameters for models over a billion parameters, making it the most parameter-efficient among methods with learnable parameters.&#x20;

In contrast, model tuning requires a separate copy of the entire model for each task, and adapter-based methods like prefix tuning and WARP involve more parameters.

#### <mark style="color:green;">Continuous vs. discrete prompts</mark>

Unlike the discrete text prompts used by GPT-3, prompt tuning learns continuous "soft prompts" through backpropagation. These soft prompts can incorporate signals from labeled examples and outperform GPT-3's few-shot learning.

#### <mark style="color:green;">Prompt location</mark>

Prompt tuning *<mark style="color:yellow;">**prepends the soft prompts to the input embeddings,**</mark>* while other methods like prefix tuning (Li and Liang, 2021) prepend prompts at every transformer layer.&#x20;

This allows prompt tuning to modify the input representations directly, letting the model update intermediate-layer task representations based on the input example.

#### <mark style="color:green;">Frozen language model</mark>

Prompt tuning keeps the pre-trained language model frozen and only tunes the soft prompts. This prevents the model from overfitting to specific datasets by memorising spurious correlations, leading to improved robustness to domain shifts compared to model tuning.

#### <mark style="color:green;">Efficient ensembling</mark>

Prompt tuning enables efficient "prompt ensembling" by learning multiple prompts for the same task while sharing the core language model parameters. This improves performance and reduces storage and inference costs compared to traditional model ensembling.

<details>

<summary><mark style="color:blue;">What is ensembling?</mark></summary>

Prompt ensembling is a technique that involves training multiple sets of soft prompts for the same task using a single frozen pre-trained language model.&#x20;

Each set of prompts can be viewed as a separate "model" that adapts the language model to the specific task.&#x20;

By combining the predictions from these multiple prompt-based models, we can create an ensemble that often outperforms individual prompt-based models and matches the performance of the best single prompt.

<mark style="color:green;">**The main advantages of prompt ensembling are:**</mark>

<mark style="color:blue;">**Improved performance:**</mark> Ensembling multiple prompts leads to better task performance compared to the average single prompt and often matches or exceeds the best individual prompt.

<mark style="color:blue;">**Parameter efficiency:**</mark> Prompt ensembling allows for the creation of multiple task-specific models while sharing the same core language model parameters. This drastically reduces storage costs compared to traditional model ensembling, where each model in the ensemble is a separate copy of the entire model.

<mark style="color:blue;">**Inference efficiency:**</mark> During inference, instead of running multiple forward passes for each model in the ensemble, we can process the input with a single forward pass using a batch size equal to the number of prompts in the ensemble. This makes inference more efficient compared to traditional model ensembling.

<mark style="color:green;">**Here's an example of how prompt ensembling works**</mark>

Let's say we have a sentiment analysis task where we need to classify movie reviews as positive or negative.&#x20;

We start by training five different sets of soft prompts (P1, P2, P3, P4, P5) on the same training data using a single frozen pre-trained language model (e.g., T5-XXL).

During inference, given a new movie review, we prepend each set of prompts to the input and run a single forward pass with a batch size of five. This gives us five different sentiment predictions, one for each prompt:

* P1: Positive
* P2: Positive
* P3: Negative
* P4: Positive
* P5: Positive

To get the final ensemble prediction, we can use a simple majority voting scheme. In this case, four out of five prompts predict "Positive," so the final ensemble prediction is "Positive."

Another example is in question answering tasks, such as SQuAD.&#x20;

We can train multiple sets of prompts on the SQuAD dataset and use them to generate multiple answers for a given question.&#x20;

The ensemble prediction can be obtained by combining the answers generated by each prompt, either by voting or by taking the answer with the highest average confidence score.

Prompt ensembling is a technique that leverages the parameter efficiency of prompt tuning to create multiple task-specific models while *<mark style="color:yellow;">**sharing the same core language model**</mark>*.&#x20;

This allows for improved performance, reduced storage costs, and efficient inference compared to traditional model ensembling.

</details>

#### <mark style="color:green;">Interpretability</mark>

Although the learned soft prompts are less interpretable than discrete text prompts, the authors find that the nearest neighbours of prompt tokens form semantic clusters, suggesting that the prompts learn "word-like" representations.

### <mark style="color:purple;">Summary</mark>

In summary, prompt tuning is a simple yet effective method for adapting large pre-trained language models to downstream tasks.&#x20;

By learning continuous soft prompts through backpropagation, prompt tuning can match the performance of model tuning while being more parameter-efficient and enabling the reuse of a single frozen model for multiple tasks.&#x20;

The effectiveness of prompt tuning increases with model scale, making it a promising approach for efficiently leveraging large language models in various applications.


# Prefix-Tuning: Optimizing Continuous Prompts for Generation

This highly cited <mark style="color:blue;">**January 2021**</mark> paper introduced a new technique for efficiently fine-tuning language models (LMs) called <mark style="color:blue;">**prefix-tuning**</mark>.

This method addressed the challenges of efficiently adapting LMs to specific tasks while maintaining their generalisation capabilities and minimising the storage requirements for task-specific parameters.

Prefix tuning adapts pre-trained language models to specific tasks without modifying the original model's weights.

Prefix tuning draws inspiration from the concept of prompting, where task instructions and examples are prepended to the input to steer the LM's generation. However, instead of using discrete tokens, <mark style="color:yellow;">**prefix tuning uses a continuous prefix vector**</mark>.

{% embed url="<https://arxiv.org/abs/2101.00190>" %}
Prefix-Tuning: Optimising Continuous Prompts for Generation" by Xiang Li
{% endembed %}

Prefix-tuning involves prepending a sequence of <mark style="color:blue;">**continuous task-specific vectors**</mark>, called a prefix, to the input of the LM.&#x20;

The Transformer can attend to these <mark style="color:blue;">**prefix vectors**</mark> as if they were a sequence of <mark style="color:yellow;">**"virtual tokens"**</mark>. Unlike prompting, the <mark style="color:blue;">**prefix vectors**</mark> do not correspond to real tokens but are learned during training.

### <mark style="color:blue;">Here's a step-by-step explanation</mark>

#### <mark style="color:green;">Soft Prompt Creation</mark>

* In prefix tuning, we create a <mark style="color:blue;">**tensor**</mark> called a "soft prompt" for each transformer block in the model.
* This soft prompt is a set of <mark style="color:blue;">**learnable parameters**</mark> that are specific to the task we want to adapt the model for.

#### <mark style="color:green;">Soft Prompt Processing</mark>

* Before using the soft prompt, it is passed through a set of <mark style="color:blue;">**fully connected layers**</mark>.
* These layers transform the soft prompt into a suitable representation that can be combined with the main input to the transformer block.

#### <mark style="color:green;">Input Modification</mark>

* The transformed soft prompt is then concatenated with the main input to the transformer block.
* This concatenation happens along the sequence length dimension, meaning the soft prompt is added as additional tokens at the beginning of the input sequence.

#### <mark style="color:green;">Transformer Block Processing</mark>

* The modified input, which now includes the soft prompt, is passed through the standard transformer block operations.
* These operations include self-attention, layer normalisation, and feed-forward neural network layers, along with residual connections.
* The transformer block processes the input as usual, but now it also takes into account the information provided by the soft prompt.

#### <mark style="color:green;">Training</mark>

* During training, only the soft prompts are updated, while the pre-trained model's weights remain frozen.
* The model learns to adapt to the specific task by adjusting the soft prompts based on the task-specific training data.
* By keeping the original model's weights unchanged, prefix tuning allows for efficient adaptation without the need for fine-tuning the entire model.

The key idea behind prefix tuning is that by adding task-specific soft prompts to each transformer block, the model can learn to condition its behavior based on the prompts.

The soft prompts act as a "prefix" that guides the model's attention and computation towards the relevant information for the task at hand.

<figure><img src="/files/L3U5SNcnUcDQ6aQeE6wi" alt="" width="563"><figcaption></figcaption></figure>

### <mark style="color:purple;">A more granular example</mark>

#### <mark style="color:green;">Define the prefix</mark>

* The prefix is a sequence of <mark style="color:blue;">**continuous vectors**</mark> that are <mark style="color:yellow;">**prepended to the input sequence**</mark>.
* The length of the prefix is a <mark style="color:blue;">**hyperparameter**</mark> that you can choose based on the complexity of the target personality and the available computational resources. Common prefix lengths range from 10 to 50 tokens.
* The prefix is initialized as a trainable matrix $$P$$ of size $$(prefixlength, embeddingdimension)$$, where $$prefixlength$$ is the <mark style="color:blue;">**number of prefix tokens**</mark> and $$embeddingdimension$$ is the <mark style="color:blue;">**size of the model's word embeddings**</mark>.
* Each row of the <mark style="color:blue;">**prefix matrix**</mark> $$P$$ corresponds to a <mark style="color:blue;">**prefix token**</mark>, and the values in that row represent the embedding of that token.
* The <mark style="color:blue;">**prefix matrix**</mark> $$P$$ is randomly initialised or initialised using the activations of real words that are relevant to the target personality. Initialising with relevant words can provide a good starting point for the prefix and potentially speed up convergence during training.

#### <mark style="color:green;">Modify the model architecture</mark>

* In the Transformer architecture, the <mark style="color:blue;">**input sequence**</mark> is typically represented as a <mark style="color:blue;">**matrix**</mark> of <mark style="color:blue;">**word embeddings**</mark>, where each row corresponds to a <mark style="color:blue;">**token in the sequence**</mark>.
* To incorporate the prefix, you concatenate the prefix matrix $$P$$ with the input embeddings matrix along the sequence dimension (usually axis 1). This results in a <mark style="color:blue;">**new input matrix**</mark> $$\[P; X]$$, where $$X$$ is the <mark style="color:blue;">**original input embeddings matrix**</mark>.
* During the <mark style="color:blue;">**forward pass**</mark>, the concatenated matrix $$\[P; X]$$ is passed through the Transformer layers, which include self-attention and feed-forward layers.
* The <mark style="color:blue;">**self-attention mechanism**</mark> in the Transformer layers allows the prefix tokens to attend to and influence the representations of the input tokens, effectively steering the model's behavior.
* Importantly, during training, only the <mark style="color:blue;">**prefix matrix**</mark> $$P$$ is updated, while the pre-trained model's parameters (i.e., the weight matrices in the Transformer layers) remain frozen. This ensures that the prefix adapts to the target personality while preserving the general language understanding captured by the pre-trained model.

#### <mark style="color:green;">Train the prefix</mark>

* To train the prefix, you use a prepared dataset that <mark style="color:yellow;">**consists of input-output pairs**</mark>, where the input is a prompt or context and the output is the corresponding response that reflects the desired personality.
* During training, you feed the <mark style="color:blue;">**input sequence**</mark> through the modified model architecture, which includes the <mark style="color:blue;">**prefix matrix**</mark> $$P$$ concatenated with the <mark style="color:blue;">**input embeddings**</mark>.
* The model generates a probability distribution over the vocabulary for each position in the output sequence, and you compute a language modeling loss (e.g., cross-entropy loss) between the predicted probabilities and the true output tokens.
* The <mark style="color:blue;">**gradients of the loss**</mark> with respect to the <mark style="color:blue;">**prefix matrix**</mark> $$P$$ are computed using <mark style="color:blue;">**backpropagation**</mark>, and the prefix matrix is updated using an optimisation algorithm like Adam.
* The pre-trained model's parameters remain fixed during training, so only the prefix matrix $$P$$ is updated to minimise the language modeling loss.
* You can experiment with different hyperparameters such as the learning rate, batch size, and number of training epochs to find the optimal configuration that achieves the best performance on a validation set.

By training the <mark style="color:blue;">**prefix matrix**</mark> $$P$$ while keeping the pre-trained model's parameters frozen, you allow the prefix to adapt to the target personality while leveraging the general language understanding captured by the pre-trained model.&#x20;

The prefix acts as a <mark style="color:yellow;">**"soft prompt"**</mark> that steers the model's behavior towards generating responses that align with the desired personality.

### <mark style="color:purple;">Prefix tuning has several advantages</mark>

* It allows for efficient adaptation of pre-trained models to new tasks without modifying the original model's weights.
* It requires fewer trainable parameters compared to fine-tuning the entire model, making it more computationally efficient.
* It can be applied to any pre-trained transformer-based model without the need for task-specific architectures.

The diagram below demonstrates:

<figure><img src="/files/XELIpo59hEJlwTcQHPEz" alt="" width="375"><figcaption><p><em>A transformer block modified for prefix tuning</em></p></figcaption></figure>

### <mark style="color:purple;">Utility Benefits of Prefix Tuning</mark>

Prefix-tuning allows for the independent training of tasks, enabling scalable personalisation without data cross-contamination. &#x20;

Each user's data can be isolated, and a *<mark style="color:yellow;">**personalised prefix can be trained for each user**</mark>*, ensuring privacy and modularity.  The independence of tasks also enables efficient batching across users and the creation of ensembles of multiple prefixes trained on the same task.

### <mark style="color:purple;">Experimental Proof</mark>

The paper demonstrates the effectiveness of prefix-tuning through extensive experiments on various natural language generation tasks, such as table-to-text generation and summarisation. &#x20;

The results show that prefix-tuning outperforms other lightweight fine-tuning methods, such as <mark style="color:blue;">**adapter-tuning**</mark>, while using substantially fewer parameters.  It achieves performance comparable to full fine-tuning, especially in low-data regimes and when generalising to unseen topics.

The authors also explore the impact of prefix length on the model's performance, revealing that there is an <mark style="color:yellow;">**optimal prefix length for each task**</mark>.

Increasing the prefix length up to a certain threshold improves performance, but further increases lead to diminishing returns and potential overfitting.

Furthermore, the paper compares prefix-tuning with an embedding-only approach, where only the embeddings of the virtual tokens are optimised.&#x20;

The results demonstrate that the <mark style="color:yellow;">embedding-only approach lacks the expressiveness necessary to achieve optimal performance</mark>, highlighting the importance of optimising the prefix vectors across all layers of the LM.

The discussion section of the paper emphasises the potential of prefix-tuning for real-world applications, particularly in scenarios requiring personalisation, privacy, efficiency, and scalability.&#x20;

The modularity and independence of tasks make <mark style="color:yellow;">prefix-tuning suitable for enterprise-level applications where customer-specific interactions and computational efficiency are crucial.</mark>

In conclusion, the paper introduces a powerful and efficient method for fine-tuning large language models.  By optimising continuous prompts in the form of prefix vectors, prefix-tuning achieves strong performance while significantly reducing the storage requirements for task-specific adaptations.&#x20;

The technique's modularity, privacy-preserving nature, and scalability make it particularly suitable for real-world applications and enterprise-level deployments.&#x20;


# Harnessing the Power of PEFT: A Smarter Approach to Fine-tuning Pre-trained Models

Parameter-Efficient Fine-Tuning (PEFT) is a technique used to fine tune neural language models

The concept of fine-tuning pre-trained models has become a cornerstone for achieving enhanced performance on specific tasks.&#x20;

However, as these models grow in complexity and size, traditional fine-tuning methods demand an increasingly hefty computational costs.

<mark style="color:blue;">Parameter-Efficient Fine-Tuning (PEFT)</mark> has been developed to optimise AI model performance efficiently, catering to scenarios where extensive retraining or large-scale parameter updates are not viable.&#x20;

Before describing the process, it is worthwhile spending some time on the architecture of neural language model.

Neural language models are built using Transformer architectures, which consist of multiple layers of self-attention and feed-forward neural networks. Each layer contains a large number of parameters, contributing to the model's ability to understand and generate complex language patterns.  When you add up multiple layers of parameters - you end up with billions of 'parameters'.

### <mark style="color:purple;">What is a parameter?</mark>

Parameters in neural language models are numerical values that define the behaviour of the model. They are the core elements that the model adjusts during the training process to learn from data.

Typically, these parameters are the weights and biases in the neural network's layers. &#x20;

<mark style="color:green;">Weights</mark> determine how much influence one node (or neuron) in a layer has on another in the subsequent layer.&#x20;

<mark style="color:green;">Biases</mark> are added to the output of weighted node inputs and provide additional flexibility to the model, allowing it to better fit the data.

These models contain billions of 'parameters' - and each one takes up memory in a computer.  Some of the larger language models require up to 600GB of memory to operate (not disk space).

### <mark style="color:purple;">Traditional Fine Tuning versus Parameter Efficient Fine Tuning</mark>

Traditional fine-tuning methods involve <mark style="color:yellow;">**updating all the parameters**</mark>  based on a specific task or dataset.   However, this approach can be resource-intensive due to the vast number of parameters  and can lead to issues like overfitting on smaller datasets.

Parameter Efficient Fine Tuning aims to *<mark style="color:yellow;">**modify only a small fraction of the model's parameters during the fine-tuning process**</mark>*. This approach seeks to retain most of the pre-trained knowledge of the model while adapting it to specific tasks or datasets, making it more efficient and resource-friendly.

{% embed url="<https://arxiv.org/abs/2312.12148>" %}
A good review of the field of Parameter Efficient Fine Tuning
{% endembed %}

### <mark style="color:purple;">Understanding Fine-tuning and PEFT</mark>

Fine-tuning involves adjusting a pre-trained model further on a new task using new data.

Traditionally, this process updates all layers and parameters of the model, requiring significant computational resources and time, particularly for larger models. This method, while effective, is not always practical or necessary for achieving optimal results on the new task.

On the flip side, *<mark style="color:yellow;">**PEFT focuses on training only a crucial subset of the model's parameters**</mark>*, significantly reducing the computation required for fine-tuning.&#x20;

By identifying and updating the most impactful parameters for the new task, PEFT offers a more resource-efficient pathway to model optimisation.

### <mark style="color:purple;">Comparative Analysis: PEFT vs. Standard Fine-tuning</mark>

<table><thead><tr><th width="231">Feature</th><th width="275">Parameter-efficient Fine-tuning</th><th>Standard Fine-tuning</th></tr></thead><tbody><tr><td><strong>Goal</strong></td><td>Improve performance on a specific task with limited data and computation</td><td>Improve performance on a specific task with ample data and computation</td></tr><tr><td><strong>Training Data</strong></td><td>Small dataset (fewer examples)</td><td>Large dataset (many examples)</td></tr><tr><td><strong>Training Time</strong></td><td>Faster, as only a subset of parameters is updated</td><td>Longer, due to updating the entire model</td></tr><tr><td><strong>Computational Resources</strong></td><td>Fewer required</td><td>Larger required</td></tr><tr><td><strong>Model Parameters</strong></td><td>Modifies only a small subset</td><td>Re-trains the entire model</td></tr><tr><td><strong>Overfitting</strong></td><td>Less prone</td><td>More prone</td></tr><tr><td><strong>Training Performance</strong></td><td>Good enough, though not as high as full fine-tuning</td><td>Typically better than PEFT</td></tr><tr><td><strong>Use Cases</strong></td><td>Ideal for low-resource settings</td><td>Ideal for high-resource settings</td></tr></tbody></table>

### <mark style="color:purple;">Benefits of PEFT Over Traditional Fine-tuning</mark>

PEFT not only streamlines the fine-tuning process but also addresses several limitations of the traditional approach:

<mark style="color:green;">**Reduced Computational and Storage Costs**</mark><mark style="color:green;">:</mark> By updating a minimal number of parameters, PEFT significantly cuts down on computational and storage demands.

<mark style="color:green;">**Overcoming Catastrophic Forgetting**</mark><mark style="color:green;">:</mark> Traditional fine-tuning risks the model forgetting previously learned information. PEFT mitigates this by limiting parameter updates.

<mark style="color:green;">**Efficiency in Low-data Regimes**</mark><mark style="color:green;">:</mark> PEFT has shown superior performance and generalisation in scenarios with limited data, making it ideal for niche applications.

<mark style="color:green;">**Portability**</mark><mark style="color:green;">:</mark> The compact nature of PEFT modifications facilitates easy deployment and application across multiple tasks, without the need to overhaul the entire model.

<mark style="color:green;">**Comparable Performance**</mark><mark style="color:green;">:</mark> Despite its efficiency, PEFT can achieve results on par with traditional fine-tuning, ensuring no compromise on model effectiveness.

### <mark style="color:purple;">PEFT: The Future of Model Optimisation</mark>

As the demand for sophisticated AI applications grows, so does the need for more efficient model training and fine-tuning methods.&#x20;

PEFT represents a significant leap forward, offering a viable solution that balances performance with computational efficiency.&#x20;

Whether you're working in a resource-constrained environment or looking to optimise a vast pre-trained model for a new task, PEFT provides a pathway to achieving high-quality results without the traditional costs.


# What is Low-Rank Adaptation (LoRA) -  explained by the inventor

Edward Hu

Edward Hu, formerly a researcher at Microsoft, introduces <mark style="color:blue;">**Low Rank Adaptation (LoRa)**</mark>, a method to efficiently customise pretrained neural networks like diffusion models or language models.&#x20;

LoRa enhances training speed and significantly reduces checkpoint sizes by <mark style="color:yellow;">**fine-tuning a minimal number of parameters**</mark>, maintaining the performance level of comprehensive fine-tuning.&#x20;

Originating in early 2021 during Microsoft's collaboration with OpenAI, LoRa addressed the limitations of few-shot prompting and the prohibitive expense of full fine-tuning, especially for models with vast numbers of parameters.

{% embed url="<https://www.youtube.com/watch?v=DhRoTONcyZE>" %}
LoRA explained by the genius inventor Edward Hu
{% endembed %}

LoRa operates by questioning the necessity and extent of parameter fine-tuning, using a 2D plane to illustrate the range of possible configurations. It controls matrix update expressivity through rank limitation, enabling substantial parameter reduction without compromising transformation capabilities.&#x20;

This method proved nearly as effective as full fine-tuning but with vastly smaller storage and quicker deployment benefits.

LoRa's applicability extends beyond language models to any architecture involving matrix multiplication, offering a clear path for adjustment if underperformance occurs.&#x20;

It significantly cuts down checkpoint sizes (e.g., reducing a 1 TB checkpoint to 25 MB) and allows for additional low-rank matrices without introducing inference latency.&#x20;

The updates are additive, allowing seamless model switching and parallel training of multiple modules. LoRa's additive nature also supports hierarchical model specialization, enabling efficient, layered fine-tuning and rapid task or user-specific model adaptation.

<figure><img src="/files/2BnLUWqjPaWpU7e3tMb8" alt=""><figcaption></figcaption></figure>

<details>

<summary><mark style="color:green;">Full Transcript</mark></summary>

Low Rank Adaptation, or LoRa, allows you to efficiently customise pre trained neural networks, such as diffusion models or language models.&#x20;

It speeds up training and drastically reduces the size of model checkpoints by training very few parameters compared to the base model, while preserving the performance of full fine tuning. And it has become one of the go to methods for customizing AI models.&#x20;

My name is Edward Hu, and I led the invention of LoRa when I was a researcher at Microsoft. In this video, I'll share the research story behind LoRa, how I understand it, and its technical benefits.

My team at Microsoft was tasked with answering, can this GPU 3 stuff actually make money? Somewhat surprising finding was that few shot prompting was not enough to get even the largest models performed well enough for production, especially for tasks like natural language to code because it rarely appears in the training data.

Fine tuning through gradient updates was a necessity. However, full fine tuning is prohibitively expensive.&#x20;

A single model checkpoint for the 175,000,000,000 parameter variant is 1 terabyte large, which is hard to store and takes minutes to load when deployed. And that's not gonna work when we need to switch among tasks and users rapidly. We tried many off the shelf, parameter efficient fine tuning methods, but all of them had a compromise for our case.

It was with product impact in mind that we invented LoRa.&#x20;

<mark style="color:green;">So what is LoRa?</mark> I like to see it as a generalization of full fine tuning by asking 2 questions.&#x20;

Question 1: Do we need to fine tune all the parameters?&#x20;

Question 2: For the weight matrices we fine tune, how expressive should the updates be in terms of matrix rank?

We can turn these two questions into the 2 axes of a 2 d plane.&#x20;

Full fine tuning is all the way in the upper right corner and the axes, including the origin, correspond to the original model. Any point in this box is a valid LoRa configuration. Let's quickly talk about how we control the expressivity of a matrix update by controlling its rank. A d by d matrix can represent any linear transformation in a d dimensional vector space.

However, if we start with a vector in Rd, first transform it to rr, where r is less than d, and finally transform it back to rd, we restrict the kind of linear transformations we can represent.&#x20;

How does stopping by rr achieve that? Imagine an extreme case where r, or rank, equals 1. Whatever the input does boils down to just one number, which can only scale the output. By picking a small r, the kind of linear transformation we can represent is greatly limited, even though the output is still in rd.

Now, we only have to store 2 times d times r parameters instead of d squared. This is how LoRa stores matrix updates. Now back to the 2 d plane we talked about.&#x20;

The surprising result of the LoRa paper is that a point near the origin performed just as well as full fine tuning all the way in the corner. Once we see LoRa as a generalization of full fine tuning, we can easily answer some commonly asked questions for using LoRa, such as how to choose the rank R, or when to use full fine tuning.

Since full fine tuning is a special case of LoRa, and we know that full fine tuning works, we can start with a point near the origin and work our way back to the corner. At some point, this has to work, and more likely than not, what ended up working will be near the origin. Otherwise, just give up and do full fine tuning. How might that happen? Let's consider a thought experiment.

If we take a language model pre trained on English and English only, but we wanna adapt it for some tasks, say in Martian. And let's say that English and Martian have little in common, since we're basically training a model all over again. Parameter efficient fine tuning methods shouldn't work too well, So we might as well do full fine tuning instead. Another question is, can I use LoRa for a certain model of architecture? Say a wave net or a support vector machine.

And by the way, nobody really asks about the latter, but as long as the model uses matrix multiplication, we can ask, do we need to fine tune all the parameters and how expressive should the updates be? As long as we can ask these 2 questions, we can use LoRa, which makes it very generally applicable. Indeed, while we invented LoRa for large language models, people later found it to be very effective for diffusion models as well. I wanna point out that one advantage of LoRa is that it's clear what to do next if underperformance. We adapt more parameters and increase the rank.

For approaches like prefix tuning, midfit, or adapter, it is not clear what we can do next because there isn't a knob to turn that allow these methods to recover full fine tuning, unlike LoRa. Now, let's dive into the specifics of the benefits of LoRa. The most visible one is a reduction of checkpoint sizes. On GP3, we reduce the checkpoint size from 1 terabyte down to 25 megabytes. This is a direct result of training much fewer parameters 4,700,000 in this case compared to 175,000,000,000.

Another important benefit is that *<mark style="color:yellow;">**LoRa doesn't introduce any inference latency.**</mark>* You might say, hold on, don't we have these additional low rank matrices on the side while that's truly trained? What happens during inference stats? Since LoRa updates are additive to the original parameters, we can expand the low rank matrices by multiplying out the low rank bottleneck and add the updates to the original parameters. Now, we can perform inference literally the same way as with a base model, and there is no additional latency by definition.

When we need to switch tasks, we simply repeat the process but this time subtract the updates. By being careful about numerical precisions, we can recover the original parameters and we repeat to load another LoRa module. This process can be done in parallel for all parameters and is faster than a single forward pass. This is how we can switch models quickly without introducing any additional inference latency. Finally, I want to mention a few engineering ideas enabled by LoRa.

The first one is to cache many LoRa modules in RAM during deployment. So model switching simply involves data transfer between RAM and VRAM. Since RAM is usually much larger than VRAM, we can cache thousands of LoRa modules and never worry about reading from the disk again. Another idea is to train multiple LoRa modules in parallel, each on its own task. This is achieved by sharing the same base model and routing different inputs in a single batch through different LoRa modules.

This way, we can batch different LoRa jobs together and fully utilize the GPUs. There are several community implementations of this, which are linked in the video description. The final idea uses the fact that LoRa models are additive. Imagine a pipeline where a pre trained model is gradually specialized, or maybe it's first fine tuned for a particular language, and then a particular domain, and finally a particular task or even a specific user. The adaptive models form a tree, Each non root node can be a Lua module on top of the sum of its ancestors.

The model rank can be larger near the root and smaller near the leaves to accommodate different dataset sizes. Model switching here becomes tree traversal and we never have to load the base model more than once. Let me know in the comments if you have any cool ideas on how LoRa can be used or extended or if you just have any questions. I'll see you in the next video.

</details>


# Low Rank Adaptation (Lora)

Fine-tuning large language models has become increasingly challenging due to the vast number of parameters involved and the computational resources required.&#x20;

As models grow in size, traditional fine-tuning methods become impractical and costly.&#x20;

<mark style="color:blue;">**Low-Rank Adaptation (LoRA)**</mark> is a solution to this problem by efficiently adapting pre-trained language models to downstream tasks while significantly reducing the number of trainable parameters.&#x20;

This document decomposes the famous <mark style="color:blue;">**October 2021**</mark> paper describing the technique.

{% embed url="<https://arxiv.org/abs/2106.09685>" %}
LoRA: Low-Rank Adaptation of Large Language Models - cited over 2,000 times
{% endembed %}

### <mark style="color:purple;">The Intrinsic Rank Hypothesis that underpins LoRA</mark> <a href="#a64b" id="a64b"></a>

In a 2020 paper called <mark style="color:blue;">**"Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning"**</mark> from the team at Facebook it was found that pre-trained language models can still learn efficiently *<mark style="color:yellow;">**even when their parameters are randomly projected onto a**</mark>*[ *<mark style="color:yellow;">**smaller subspace**</mark>*](#user-content-fn-1)[^1].&#x20;

The intrinsic rank hypothesis is a crucial concept that underlies the development of LoRA. &#x20;

The hypothesis suggests that the model's parameters lie in a lower-dimensional subspace than the full parameter space, referred to as the <mark style="color:blue;">**"intrinsic dimension"**</mark> of the model.

intrinsic dimension represents the minimum number of dimensions required to accurately capture the important information in the model's parameter space.

The intrinsic rank hypothesis *<mark style="color:yellow;">**extends this idea to weight updates that occur during fine-tuning of language models**</mark>*.&#x20;

It posits that the updates to the weights also have a low <mark style="color:blue;">**"intrinsic rank"**</mark>**,** meaning they can be well-approximated by a <mark style="color:blue;">**low-rank matrix**</mark>.  &#x20;

This hypothesis is significant because it implies that the necessary adjustments to the model during fine-tuning can be captured using a lower-dimensional representation, *<mark style="color:yellow;">**rather than updating the entire weight matrix**</mark>*.

By exploiting the <mark style="color:blue;">**intrinsic low rank**</mark> of the <mark style="color:blue;">**weight updates**</mark>, LoRA can significantly reduce the number of trainable parameters while still allowing for effective adaptation to downstream model tasks.&#x20;

This not only makes fine-tuning more computationally efficient but also enables the adaptation of large language models on resource-constrained devices.&#x20;

The <mark style="color:blue;">**intrinsic rank hypothesis**</mark> is a fundamental principle that guides the design and implementation of LoRA.

<details>

<summary><mark style="color:green;"><strong>What are low rank matrices?</strong></mark></summary>

A <mark style="color:blue;">**matrix**</mark> is a rectangular array of numbers arranged in <mark style="color:blue;">**rows and columns**</mark>.&#x20;

The <mark style="color:blue;">**rank of a matrix**</mark> is the *<mark style="color:yellow;">**maximum number of linearly independent rows or columns**</mark>* in the matrix.&#x20;

In other words, it's the dimension of the vector space spanned by the matrix's rows or columns.

Consider a <mark style="color:blue;">**matrix**</mark> A:

```yaml
A = [1 2 3]
    [2 4 6]
    [3 6 9]
```

In this matrix, we can see that the second row is 2 times the first row, and the third row is 3 times the first row.&#x20;

This means that the rows are <mark style="color:blue;">**linearly dependent**</mark>. We can express any row as a linear combination of the other rows.

Similarly, the second column is 2 times the first column, and the third column is 3 times the first column. The columns are also <mark style="color:blue;">**linearly dependent**</mark>.

In this case, the rank of the matrix is 1. Despite the matrix being 3x3, it only contains one independent piece of information.

Now, let's look at the *<mark style="color:yellow;">**concept of lower-rank matrices**</mark>.*&#x20;

A matrix is considered to be of lower rank if its <mark style="color:yellow;">**rank is less than the minimum of its number of rows and columns**</mark>. In the example above, the matrix has a rank of 1, which is lower than min(3, 3) = 3, so it is a lower-rank matrix.

The idea of lower-rank matrices is used in many applications, such as:

<mark style="color:blue;">**Data Compression**</mark>: By approximating a matrix with a lower-rank matrix, we can store less data while preserving the most important information.

<mark style="color:blue;">**Recommendation Systems:**</mark> User-item matrices in recommendation systems are often of lower rank because user preferences can be described by a smaller number of latent factors.

<mark style="color:blue;">**Image Processing:**</mark> Many operations in image processing, such as image denoising and compression, exploit the fact that image matrices are often of lower rank.

The <mark style="color:blue;">**Rank-Nullity Theorem**</mark> states that for a linear map (which can be represented by a matrix) between two vector spaces, the dimension of the domain (number of columns) equals the sum of the rank (dimension of the image) and the nullity (dimension of the kernel).&#x20;

This theorem connects the concepts of rank and nullity, showing that they are complementary.

In summary, lower-rank matrices are matrices whose rank is less than the maximum possible given their dimensions.&#x20;

They are used in many applications to simplify data, reduce dimensionality, and uncover hidden structures. The rank of a matrix can be determined by the number of linearly independent rows or columns, which are always equal, as stated by the Rank-Nullity Theorem.

</details>

### <mark style="color:purple;">**Weight Matrices in Transformers**</mark>

In the Transformer architecture, the <mark style="color:blue;">**self-attention layer**</mark> is a key component that allows the model to attend to different positions of the <mark style="color:blue;">**input sequence**</mark>.&#x20;

The <mark style="color:blue;">**self-attention mechanism**</mark> is applied to the <mark style="color:blue;">**input embeddings**</mark> or the output of the previous layer, which we'll denote as $$X$$, with <mark style="color:blue;">**shape**</mark> $$(sequencelength, dmodel)$$.

The self-attention layer consists of <mark style="color:blue;">**multiple attention heads that operate in parallel**</mark>.&#x20;

Each <mark style="color:blue;">**attention head**</mark> performs the following steps:

1. Linearly project the <mark style="color:blue;">**input**</mark> $$X$$ into <mark style="color:blue;">**query**</mark>, <mark style="color:blue;">**key**</mark>, and <mark style="color:blue;">**value**</mark> representations using the corresponding <mark style="color:blue;">**weight matrices**</mark> ( $$Wq$$, $$Wk$$, $$Wv$$).
2. Compute the <mark style="color:blue;">**attention scores**</mark> by taking the <mark style="color:blue;">**dot product**</mark> of the <mark style="color:blue;">**query**</mark> and <mark style="color:blue;">**key**</mark> representations.
3. Scale the attention scores and apply a <mark style="color:blue;">**softmax function**</mark> to obtain the <mark style="color:blue;">**attention weights**</mark>.
4. <mark style="color:blue;">**Multiply the attention weights**</mark> with the <mark style="color:blue;">**value representations**</mark> to get the <mark style="color:blue;">**weighted values**</mark>.
5. <mark style="color:blue;">**Concatenate the weighted values**</mark> from all attention heads and linearly project them using the <mark style="color:blue;">**output projection matrix**</mark> $$(Wo)$$.

In the Transformer architecture, there are <mark style="color:yellow;">**four**</mark>**&#x20;**<mark style="color:blue;">**weight matrices**</mark>**&#x20;in** the self-attention module:&#x20;

Now, let's focus on the <mark style="color:yellow;">**dimensions of the weight matrices**</mark>:

* <mark style="color:blue;">**Query matrix**</mark> $$(Wq)$$: $$(dmodel, d\_q)$$
* <mark style="color:blue;">**Key matrix**</mark> $$(Wk)$$: $$(dmodel, d\_k)$$
* <mark style="color:blue;">**Value matrix**</mark> $$(Wv)$$: $$(dmodel, d\_v)$$
* <mark style="color:blue;">**Output projection matrix**</mark> $$(Wo)$$: $$(dmodel, dmodel)$$

<mark style="color:blue;">**Weight matrices**</mark> enable the Transformer model to learn and capture complex relationships and dependencies within the input sequence.  These learned weights in these matrices are crucial for the model's ability to attend to relevant information and generate meaningful representations.

In summary, the self-attention mechanism in the Transformer architecture relies on four key weight matrices: query, key, value, and output projection matrices.&#x20;

These matrices linearly project the input into different subspaces, enabling the model to capture dependencies and learn meaningful representations.  The dimensions of these matrices are typically determined by the model's hyperparameters, such as the dimensionality of the model and the number of attention heads.

Understanding the roles and dimensions of these weight matrices is crucial for grasping the concepts behind LoRA and the Intrinsic Rank Hypothesis, which we will explore in the next section.&#x20;

### <mark style="color:purple;">How does LoRA work?</mark>

LoRA (Low-Rank Adaptation) introduces a *<mark style="color:yellow;">**modification to the weight matrices**</mark>* to efficiently adapt the pre-trained model to downstream tasks.&#x20;

Let's see how LoRA gets involved in the process and influences the weights.

Recall the <mark style="color:blue;">**self-attention mechanism**</mark> has four <mark style="color:blue;">**weight matrices**</mark>:

* <mark style="color:blue;">**Query matrix**</mark> $$(Wq)$$
* <mark style="color:blue;">**Key matrix**</mark> $$(Wk)$$
* <mark style="color:blue;">**Value matrix**</mark> $$(Wv)$$
* <mark style="color:blue;">**Output projection matrix**</mark> $$(Wo)$$

These matrices are typically learned during the pre-training phase and have <mark style="color:blue;">**full rank**</mark>.

&#x20;In linear algebra, a *<mark style="color:yellow;">matrix is said to have</mark> <mark style="color:yellow;"></mark><mark style="color:yellow;">**full rank**</mark> <mark style="color:yellow;"></mark><mark style="color:yellow;">if its rank is equal to the smaller of its number of rows or columns</mark>*.&#x20;

In the context of neural networks, this means that the <mark style="color:blue;">**weight matrices**</mark> in dense layers are typically not low-rank - they cannot be exactly represented as a product of two smaller matrices.

LoRA <mark style="color:yellow;">**modifies the weight matrices**</mark> by introducing a low-rank decomposition of the weight updates.&#x20;

Instead of directly updating the pre-trained weight matrices, LoRA represents the weight updates *<mark style="color:yellow;">**using two smaller matrices**</mark>*, $$A$$ and $$B$$, such that:

$$Wupdated = Wpretrained + ∆W$$

&#x20;$$∆W = B \* A$$

where:

* $$Wpretrained$$ is the <mark style="color:blue;">**original pre-trained weight matrix**</mark> ($$Wq, Wk, Wv, Wo$$)
* $$∆W$$ is the <mark style="color:blue;">**weight update matrix**</mark>
* $$B$$ is a <mark style="color:blue;">**matrix**</mark> of size $$(dmodel, r$$), where $$r$$ is the rank of the decomposition
* $$A$$ is a <mark style="color:blue;">**matrix**</mark> of size $$(r, dmodel)$$

The key idea behind LoRA is to use a low-rank decomposition of the weight updates.  By representing the weight updates as the product of two smaller matrices, A and B, LoRA allows for a more compact and efficient representation.&#x20;

The <mark style="color:blue;">**rank (r)**</mark> of the decomposition determines the size of these matrices and controls the expressiveness of the adaptation. &#x20;

A lower rank results in fewer trainable parameters, making the adaptation more memory-efficient, while a higher rank allows for more flexibility in adapting the weights.

### <mark style="color:purple;">During the</mark> <mark style="color:purple;"></mark><mark style="color:purple;">**fine-tuning process**</mark> <mark style="color:purple;"></mark><mark style="color:purple;">with LoRA</mark>

The method involves freezing the original model weights and <mark style="color:yellow;">**adjusting only two smaller matrices**</mark>, $$A$$ <mark style="color:blue;">**and**</mark> $$B$$.

1. The <mark style="color:blue;">**pre-trained weight matrices**</mark> $$(Wpretrained)$$ <mark style="color:yellow;">**remain frozen**</mark> and do not receive gradient updates.
2. The <mark style="color:blue;">**matrices**</mark> $$A$$ and $$B$$ are initialised randomly and *<mark style="color:yellow;">**are the only trainable parameters**</mark>*. Matrix $$A$$ is initialised with a random Gaussian distribution, while matrix $$B$$is initialised with zeros.

<details>

<summary><mark style="color:green;"><strong>Gaussian initialisation</strong></mark></summary>

In the LoRA method, matrix A is initialised with a random Gaussian distribution, while matrix B is initialised with zeros.&#x20;

The reason for using Gaussian initialisation for matrix A is to *<mark style="color:yellow;">**introduce randomness and break symmetry**</mark>* in the initial values of the weights.

When the weights of a neural network are initialised to the same value (e.g., all zeros), the network *<mark style="color:yellow;">**may struggle to learn meaningful patterns because all the neurons behave identically**</mark>*.

By initialising the weights with random values drawn from a Gaussian distribution, we ensure that the neurons start with different initial activations, allowing them to learn diverse features during training.

The choice of Gaussian initialisation is based on the principle of <mark style="color:blue;">**"symmetry breaking"**</mark> and the idea that the weights should be initialised with <mark style="color:blue;">**small random values**</mark> to facilitate gradient flow and prevent vanishing or exploding gradients. Gaussian initialisation has been shown to work well in practice and is commonly used in deep learning.

In the context of LoRA, initialising matrix $$A$$ with a Gaussian distribution ensures that the initial weight update matrix $$∆W$$ (which is the product of $$A$$ and $$B$$) has random, small values. &#x20;

This allows the model to *<mark style="color:yellow;">**gradually adapt the pre-trained weights to the downstream task during fine-tuning**</mark>*, starting from a point of random perturbation.

By initialising matrix $$B$$ with zeros, the initial weight update matrix $$∆W$$ is effectively zero, meaning that the model starts with the original pre-trained weights.&#x20;

As training progresses, the values of $$A$$ and $$B$$ are updated based on the gradients, allowing the model to learn the necessary adaptations for the specific task.

The combination of Gaussian initialisation for matrix of $$A$$and zero initialisation for matrix of $$B$$ in LoRA ensures a balanced starting point for fine-tuning, facilitating the learning of task-specific adaptations while leveraging the knowledge captured in the pre-trained weights.

</details>

1. In the <mark style="color:blue;">**forward pass**</mark>, the input is multiplied with both the <mark style="color:blue;">**pre-trained weight matrix**</mark> $$(Wpretrained)$$ and the <mark style="color:blue;">**LoRA weight update matrix**</mark> $$(∆W = B \* A)$$. The results are then <mark style="color:blue;">**summed element-wise**</mark> to obtain the updated output.
2. During <mark style="color:blue;">**backpropagation**</mark>, gradients are computed only for the $$A$$ and $$B$$  <mark style="color:blue;">**matrices**</mark>, while the <mark style="color:blue;">**pre-trained weight matrices**</mark> remain unchanged.
3. The optimisation process updates the $$A$$ and $$B$$ <mark style="color:blue;">**matrices**</mark> based on the <mark style="color:blue;">**gradients**</mark>, allowing the model to adapt to the downstream task.

<details>

<summary><mark style="color:green;">Forward Pass and Backpropagation explanation</mark></summary>

<mark style="color:blue;">**Forward Pass**</mark>

During the forward pass with LoRA, the input is multiplied by both the pre-trained weight matrix $$Wpretrained$$ and the weight update matrix $$∆W$$. The updated output is computed as follows:

$$output = input × Wpretrained + input × ∆W ∆W = B × A$$

The pre-trained weight matrix $$Wpretrained$$ remains frozen, while the matrices $$A$$ and $$B$$ are learned during fine-tuning.&#x20;

The output of the self-attention layer is obtained by summing the results of the matrix multiplications.

<mark style="color:blue;">**Backpropagation**</mark>

During backpropagation, the *<mark style="color:yellow;">**gradients are computed with respect to the input and the trainable parameters**</mark>*.&#x20;

In LoRA, only the matrices matrices $$A$$ and $$B$$ are updated based on the gradients, while the pre-trained weight matrix $$Wpretrained$$ remains unchanged.

The gradients of the loss with respect to $$A$$ and $$B$$  are computed using the chain rule:

$$∂loss / ∂A = (∂loss / ∂∆W) × B^T ∂loss / ∂B = (∂loss / ∂∆W)^T × A$$

The optimiser then uses these gradients to update the values of $$A$$ and $$B$$

</details>

By representing the weight updates using a <mark style="color:blue;">**low-rank decomposition**</mark> $$(∆W = B \* A)$$, LoRA significantly reduces the number of trainable parameters.&#x20;

The <mark style="color:blue;">**rank r**</mark> <mark style="color:yellow;">**determines the**</mark><mark style="color:yellow;">**&#x20;**</mark>*<mark style="color:yellow;">**size**</mark>*<mark style="color:yellow;">**&#x20;**</mark><mark style="color:yellow;">**of the matrices**</mark> $$A$$ and $$B$$ and *<mark style="color:yellow;">**controls the expressiveness**</mark>* of the adaptation. &#x20;

A smaller <mark style="color:blue;">**rank r**</mark> results in fewer trainable parameters and more efficient adaptation, while a larger <mark style="color:blue;">**rank r**</mark> allows for more flexibility in adapting the weights.

### <mark style="color:purple;">What are the low rank matrices?</mark>

In the Low-Rank Adaptation (LoRA) method proposed in this paper, the terms $$A$$ and $$B$$ refer to the <mark style="color:blue;">**low-rank matrices**</mark> used to approximate the <mark style="color:blue;">**weight update matrix**</mark> $$∆W$$during adaptation.&#x20;

As discussed, the <mark style="color:blue;">**weight matrix**</mark> $$W$$ being targeted is *<mark style="color:yellow;">**part of the Transformer architecture**</mark>*, specifically the weight matrices in the self-attention module.

### <mark style="color:purple;">Matrices ( A ) and ( B )</mark> <a href="#id-230a" id="id-230a"></a>

As highlighted. the authors of Lora hypothesised that during fine-tuning, the updates to the weights $$(∆W)$$ *<mark style="color:yellow;">**have a low "intrinsic rank"**</mark>*, meaning they can be well-approximated by a <mark style="color:blue;">**low-rank matrix**</mark>.

This means that significant changes to the neural network can be captured <mark style="color:yellow;">**using a lower-dimensional representation**</mark>.&#x20;

Essentially, the idea is that <mark style="color:yellow;">**not all elements of**</mark>**&#x20;(**$$Δ W$$ <mark style="color:yellow;">**) are equally important**</mark>; instead, a smaller subset of these changes can effectively encapsulate the necessary adjustments.

Building on this hypothesis, LoRA proposes representing ($$Δ W$$ ) as the <mark style="color:blue;">**product**</mark> of two smaller matrices, $$( A )$$and $$( B )$$, *<mark style="color:yellow;">**with a lower rank**</mark>*. &#x20;

$$Δ W$$ denotes the *<mark style="color:yellow;">**relative change**</mark>* to the initial value when it has been trained.

The <mark style="color:blue;">**updated weight matrix**</mark> $$( W’ )$$ thus becomes:

$$\[ W’ = W + BA ]$$

In this equation, $$( W )$$remains frozen (i.e., <mark style="color:yellow;">**it is not updated during training**</mark>).&#x20;

The <mark style="color:blue;">**matrices**</mark> $$( B )$$ and $$( A )$$are of lower dimensionality, with their <mark style="color:blue;">**product**</mark> $$( BA )$$representing a low-rank approximation of $$( Δ W )$$.

### <mark style="color:purple;">Impact of Lower Rank on Trainable Parameters</mark> <a href="#id-1131" id="id-1131"></a>

By choosing <mark style="color:blue;">**matrices**</mark> $$( A )$$ and $$( B )$$ to have a <mark style="color:blue;">**lower rank**</mark> $$( r )$$, the number of trainable parameters is significantly reduced.&#x20;

For example, if $$( W )$$ is a $$( d x d )$$ <mark style="color:blue;">**matrix**</mark>, traditionally, updating $$( W )$$ would involve $$( d² )$$ <mark style="color:blue;">**parameters**</mark>.&#x20;

However, with  $$( B )$$ and  $$( A )$$  of sizes $$( d  X  r )$$and $$( r Xd )$$ respectively, the total number of <mark style="color:blue;">**parameters**</mark> reduces to $$( 2dr )$$, which is much smaller when $$( r << d ).$$

The reduction in the number of trainable parameters, as achieved through the Low-Rank Adaptation (LoRA) method, offers several significant benefits, particularly when fine-tuning large-scale neural networks:

### <mark style="color:purple;">What does 'rank' mean?</mark>

The term "rank" in the context of LoRA refers to the <mark style="color:blue;">**rank of the weight update matrix**</mark> $$∆W$$, which is approximated by the <mark style="color:blue;">**product**</mark> of two smaller <mark style="color:blue;">**matrices**</mark> $$A$$ and $$B$$.&#x20;

The idea here is to exploit the fact that matrices can contain “duplicate information” in the form of linear dependence.  &#x20;

This means we can use factorisation to <mark style="color:yellow;">**represent a large matrix in terms of two smaller matrices**</mark>.

This is similar to how a large number can be represented as the multiplication of two smaller numbers, a matrix can be thought of as the multiplication of two smaller matrices.

The <mark style="color:blue;">**rank**</mark> of a <mark style="color:blue;">**matrix**</mark> is the maximum *<mark style="color:yellow;">**number of linearly independent rows or columns in the matrix**</mark>*.

In other words, it is the dimension of the vector space spanned by the rows or columns of the matrix.

<mark style="color:green;">**Example:**</mark> Consider the following <mark style="color:blue;">**matrix**</mark> M:

```yaml
[1 2 3] 
[2 4 6] 
[3 6 9]
```

The rows of this matrix are <mark style="color:blue;">**not linearly independent**</mark> because the third row is a linear combination of the first two rows $$(3 \* row1 = row3)$$.&#x20;

Similarly, the <mark style="color:blue;">**columns are not linearly independent**</mark> because the third column is a linear combination of the first two columns $$(3 \* column1 = column3).$$

To find the rank of the matrix, we can use <mark style="color:blue;">**Gaussian elimination**</mark> to convert the matrix into <mark style="color:blue;">**row echelon form**</mark>.&#x20;

<mark style="color:blue;">**Row echelon form**</mark> is a type of matrix in which:

1. All nonzero rows are above any rows of all zeros.
2. The leading entry of each nonzero row after the first occurs to the right of the leading entry of the previous row.
3. The leading entry in any nonzero row is 1.
4. All entries in the column below a leading 1 are zeros.

Using <mark style="color:blue;">**Gaussian elimination**</mark>, a matrix is transformed into row echelon form. This helps determine properties such as the rank and solve linear systems.

<mark style="color:blue;">**Gaussian elimination**</mark> is a method used to <mark style="color:yellow;">**solve systems of linear equations**</mark>, find the <mark style="color:blue;">**rank r**</mark>, and calculate the determinant of a matrix. &#x20;

The process involves three main steps:

1. <mark style="color:purple;">**Forward Elimination:**</mark> Transform the matrix into an upper triangular form.
2. <mark style="color:purple;">**Pivoting:**</mark> Swap rows to position the highest absolute value as the pivot element to avoid division by zero and improve numerical stability.
3. <mark style="color:purple;">**Back Substitution:**</mark> Solve for the variables starting from the last row upwards.

By converting the <mark style="color:blue;">**matrix**</mark> into row <mark style="color:blue;">**echelon form**</mark>, Gaussian elimination <mark style="color:yellow;">**simplifies the system**</mark>, making it easier to understand its properties and solutions.

After performing Gaussian elimination, we get <mark style="color:blue;">**matrix M**</mark>:

```
[1 2 3]
[0 1 1]
[0 0 0]
```

The number of non-zero rows in the <mark style="color:blue;">**row echelon form**</mark> is the rank of the <mark style="color:blue;">**matrix**</mark>.

In this case, the rank is 2, which is the maximum number of linearly independent rows or columns in the <mark style="color:blue;">**matrix**</mark> M.

In the context of LoRA, the <mark style="color:blue;">**rank r**</mark> determines the dimensionality of the subspace in which the <mark style="color:blue;">**weight update matrix**</mark> $$∆W$$ is approximated.&#x20;

By choosing a lower <mark style="color:blue;">**rank r**</mark>, we enforce a low-dimensional structure on the weight update matrix, reducing the number of trainable parameters while still capturing the most important aspects of the adaptation.

So to sum up, in LoRA, the <mark style="color:blue;">**rank r**</mark> is the hyperparameter that determines the <mark style="color:blue;">**size of the matrices A and B**</mark>.&#x20;

<figure><img src="/files/Y4vwNhrWF3svT0HFUsF1" alt=""><figcaption><p>A conceptual diagram of LoRA with an r value equal to 1 and 2. In both examples the decomposed A and B matrices result in the same sized change matrix, but r=2 is able to encode more linearly independent information into the change matrix, due to having more information in the A and B matrices.  Source: <a href="https://medium.com/@danielwarfield1?source=post_page-----e944a6bff46b--------------------------------">Daniel Warfield</a></p></figcaption></figure>

<mark style="color:purple;">**Specifically:**</mark>

* Matrix $$A$$ has shape $$(r, dmodel)$$, where r is the <mark style="color:blue;">**rank of the decomposition**</mark> and $$dmodel$$ is the <mark style="color:blue;">**dimension of the model**</mark>.
* Matrix $$B$$ has <mark style="color:blue;">**dimensions**</mark> $$(dmodel, r)$$

So, the updated equation is:

$$Wupdated = Wpretrained + ∆W$$&#x20;

$$∆W = B \* A$$

<mark style="color:purple;">**where:**</mark>

* $$Wpretrained$$ is the original <mark style="color:blue;">**pre-trained weight matrix**</mark> $$(dmodel, dmodel)$$
* $$∆W$$is the <mark style="color:blue;">**weight update matrix**</mark> $$(dmodel, dmodel)$$
* $$B$$ is a <mark style="color:blue;">**matrix of size**</mark> $$(dmodel, r)$$
* $$A$$ is a <mark style="color:blue;">**matrix of size**</mark> $$(r, dmodel)$$

The <mark style="color:blue;">**product**</mark> of $$A$$ and $$B$$ results in a <mark style="color:blue;">**matrix**</mark> $$∆W$$of <mark style="color:blue;">**shape**</mark> $$(dmodel, dmodel)$$, which has <mark style="color:blue;">**rank r**</mark>.&#x20;

By choosing a smaller value for r, we enforce a low-rank structure on the <mark style="color:blue;">**weight update matrix**</mark> $$∆W$$.

#### <mark style="color:green;">**Expressiveness and Rank**</mark>

The <mark style="color:blue;">**rank r**</mark> of the <mark style="color:blue;">**weight update matrix**</mark> ∆W controls the <mark style="color:blue;">**expressiveness**</mark> of the adaptation.&#x20;

A *<mark style="color:yellow;">**higher rank allows for more flexibility in adapting the weights**</mark>*, as it can capture more complex patterns and transformations.  However, increasing the rank also increases the number of trainable parameters.

On the other hand, a *<mark style="color:yellow;">**lower rank restricts the expressiveness of the adaptation**</mark>* but results in fewer trainable parameters.&#x20;

This is because the  <mark style="color:blue;">**matrices**</mark> $$( A )$$ and $$( B )$$ have fewer elements when <mark style="color:blue;">**r**</mark> is smaller.

A low-rank approximation of the <mark style="color:blue;">**weight update matrix**</mark> can still capture the most important aspects of the adaptation while being more parameter-efficient.

The choice of <mark style="color:blue;">**rank r**</mark> depends on the complexity of the downstream task and the available computational resources. It is a trade-off between expressiveness and efficiency.

The consensus is that when the data is similar to the data used in pre-training, a low <mark style="color:blue;">**rank r**</mark> value is probably sufficient.  When fine tuning on very new tasks, which might require substantial logical changes within the model, a high <mark style="color:blue;">**rank r**</mark> value may work better.

### <mark style="color:purple;">**Applying LoRA to Transformer Self Attention Weights**</mark>

* LoRA <mark style="color:yellow;">**can be applied to any or all of the self-attention weight matrices**</mark> $$(Wq, Wk, Wv, Wo)$$ in each Transformer layer.
* The paper primarily focuses on adapting only the <mark style="color:blue;">**query**</mark> $$(Wq)$$ and <mark style="color:blue;">**value**</mark> $$(Wv)$$ matrices because they play the most critical role in capturing and transforming the input representations.&#x20;
* During the <mark style="color:blue;">**forward pass**</mark>, the <mark style="color:blue;">**adapted weight matrix**</mark> $$W'$$ is computed as $$W' = W + BA$$, where $$W$$ is the <mark style="color:blue;">**pre-trained weight matrix**</mark> and $$'$$ is the low-rank adaptation term.
* This allows for efficient adaptation *<mark style="color:yellow;">**without introducing additional inference latency**</mark>*, as the low-rank matrices can be merged with the pre-trained weights during deployment.

The key idea behind LoRA is to <mark style="color:yellow;">**exploit the low-rank structure of the adaptation matrix**</mark> $$∆W$$.&#x20;

By approximating $$∆W$$ with <mark style="color:blue;">**low-rank matrices**</mark> $$A$$ and $$B$$, LoRA significantly reduces the number of trainable parameters while still allowing for effective adaptation to downstream tasks.&#x20;

This makes it suitable for adapting large pre-trained language models, where standard fine-tuning would be prohibitively expensive in terms of storage and computation.

### <mark style="color:purple;">The problem with full fine tuning</mark>

To remind us of the problem LoRA is solving, below is a brief description of the language modeling problem for a full fine tuning. In particular, the maximisation of conditional probabilities given a task-specific prompt.

* $$PΦ(y|x)$$ represents a <mark style="color:blue;">**pre-trained autoregressive language model**</mark>, where Φ denotes the <mark style="color:blue;">**model parameters**</mark>.
* $$Z = {(xi, yi)}i=1,..,N$$ is a <mark style="color:blue;">**dataset of context-target pairs**</mark> for a downstream task, where $$xi$$ is the <mark style="color:blue;">**context**</mark> and $$yi$$ is the <mark style="color:blue;">**target sequence**</mark>.
* $$Φ0$$represents the initial <mark style="color:blue;">**pre-trained weights**</mark> of the model.
* $$∆Φ$$ represents the <mark style="color:blue;">**task-specific parameter**</mark> increment during fine-tuning.
* $$Θ$$is a <mark style="color:blue;">**smaller set of parameters**</mark> used to encode $$∆Φ$$.

### <mark style="color:green;">**Language Modeling Objective**</mark>

#### The goal is to adapt a pre-trained language model to downstream conditional text generation tasks.   &#x20;

#### The model learns to <mark style="color:yellow;">**maximise the conditional probability of the target sequence**</mark> given the context.

During <mark style="color:blue;">**full fine-tuning**</mark>, the objective is to maximise the sum of log probabilities of each <mark style="color:blue;">**token**</mark> $$yt$$ in the <mark style="color:blue;">**target sequence**</mark> $$y$$, conditioned on the <mark style="color:blue;">**context**</mark> $$x$$ and the <mark style="color:blue;">**previous tokens**</mark> $$y\<t$$.&#x20;

This is achieved by updating the <mark style="color:blue;">**model parameters**</mark> from $$Φ0$$ to $$Φ0 + ∆Φ$$through [<mark style="color:blue;">**gradient descent**</mark>](#user-content-fn-2)[^2]:

<figure><img src="/files/2RtjgGbdATG171raYl1y" alt=""><figcaption></figcaption></figure>

The notation $$(x, y) ∈ Z$$ indicates that the <mark style="color:yellow;">**summation is performed over all context-target pairs**</mark> in the dataset $$Z.$$   $$|y|$$denotes the <mark style="color:blue;">**length of the target sequence**</mark> $$y$$.

#### <mark style="color:green;">**Parameter-Efficient Approach**</mark>

The main drawback of full fine-tuning is that for each downstream task, a <mark style="color:blue;">**separate set of parameters**</mark> ∆Φ is learned, which <mark style="color:yellow;">**has the**</mark><mark style="color:yellow;">**&#x20;**</mark>*<mark style="color:yellow;">**same dimension**</mark>*<mark style="color:yellow;">**&#x20;**</mark><mark style="color:yellow;">**as the pre-trained weights**</mark> Φ0.&#x20;

This can be challenging to store and deploy, especially for large models.

To address this, the authors propose a parameter-efficient approach where the <mark style="color:yellow;">**task-specific parameter increment**</mark> ∆Φ is encoded by a much <mark style="color:yellow;">**smaller set of parameters**</mark> Θ, such that |Θ| << |Φ0|.

The objective becomes:

<figure><img src="/files/Fmn9vVmion6JbwSGU12a" alt=""><figcaption></figcaption></figure>

Here, ∆Φ is a function of Θ, denoted as ∆Φ(Θ).&#x20;

The goal is to find the optimal Θ that maximises the <mark style="color:yellow;">**conditional language modeling objective**</mark>.

### <mark style="color:green;">**Low-Rank Representation**</mark>

The authors propose to use a <mark style="color:blue;">**low-rank representation**</mark> to encode $$∆Φ$$, which is both compute- and memory-efficient.&#x20;

This means that $$∆Φ$$ is represented as a <mark style="color:blue;">**product of smaller matrices**</mark>, reducing the number of parameters needed to store and update during fine-tuning.

The key idea is to <mark style="color:yellow;">**significantly reduce the number of trainable parameters**</mark> $$|Θ|$$ compared to the size of the pre-trained weights $$|Φ0|$$.

For example, when using GPT-3 175 billion parameter model as the pre-trained model, the <mark style="color:yellow;">**number of trainable parameters**</mark> $$|Θ|$$can be as *<mark style="color:yellow;">**small as 0.01%**</mark>* of $$|Φ0|$$, greatly reducing the storage and computational requirements for fine-tuning on downstream tasks.

### <mark style="color:purple;">Issues with existing solutions</mark>

While there have been other <mark style="color:blue;">**Parameter Efficient Fine Tuning (PEFT)**</mark> solutions for efficient model adaptation in transfer learning, such as <mark style="color:blue;">**adding adapter layers**</mark> or <mark style="color:blue;">**optimising input layer activations,**</mark> have limitations, especially in large-scale and latency-sensitive production scenarios.  &#x20;

<mark style="color:purple;">**We discuss below:**</mark>

### <mark style="color:green;">Adapter Layers</mark>

Adapter layers, as proposed by Houlsby et al. (2019) and Lin et al. (2020), are <mark style="color:yellow;">**additional layers inserted into the Transformer architecture**</mark> to enable parameter-efficient fine-tuning.&#x20;

While adapters have fewer parameters compared to the original model, they <mark style="color:yellow;">**introduce extra computation**</mark> that must be processed sequentially, leading to increased inference latency.

The latency issue becomes more prominent in online inference settings where the batch size is small (for example a batch size of 1).&#x20;

Even with a small bottleneck dimension in the adapter layers, the added latency can be significant (e.g., 20-30% slower for GPT-2 medium on a single GPU).  This problem is amplified when using model parallelism techniques, as the additional depth requires more synchronous GPU operations.

### <mark style="color:green;">**Optimising Input Layer Activations (Prompt Tuning)**</mark>

&#x20;Another PEFT approach is prefix tuning (Li & Liang, 2021), which *<mark style="color:yellow;">**directly optimises a portion of the input layer activations**</mark>* (the prompt) while keeping the pre-trained model unchanged.&#x20;

However, this method faces optimisation challenges and exhibits <mark style="color:blue;">**non-monotonic performance**</mark> changes with respect to the number of trainable parameters.

#### <mark style="color:blue;">Non-Monotonic Performance Changes in Prompt Tuning</mark>

<mark style="color:blue;">**Non-monotonic performance**</mark> changes refer to <mark style="color:yellow;">**fluctuations in model performance**</mark> that do not consistently improve or degrade as the number of trainable parameters increases.

In the context of prompt tuning, this means that *<mark style="color:yellow;">**increasing the number of trainable parameters does not guarantee a corresponding increase in model performance**</mark>*.&#x20;

Instead, performance may improve up to a point, then degrade, and potentially improve again, creating a non-linear relationship. These irregular performance patterns can complicate the optimisation process and make it challenging to determine the optimal parameter configuration for prompt tuning.

### <mark style="color:purple;">Lora in Practice</mark>

#### <mark style="color:green;">**What subset of weight matrices should be adopted for maximum downstream performance?**</mark>

The authors experimented with applying LoRA to different subsets of the <mark style="color:blue;">**self-attention weight matrices**</mark> when training GPT3, with a budget of 18 million parameters (this compares to GPT-3's 175 billion parameters!  The <mark style="color:blue;">**weight matrices**</mark> are as follows:&#x20;

$$Wq (query)$$

$$Wk (key)$$

$$Wv (value)$$

$$Wo (output)$$

They found that adapting both the <mark style="color:blue;">**query weights**</mark> $$Wq$$and <mark style="color:blue;">**value weights**</mark> $$Wv$$ yielded the best performance on downstream tasks like WikiSQL[^3] and MultiNLI[^4].&#x20;

Adapting $$Wq$$, $$Wk$$, $$Wv$$, and $$Wo$$ together also performed well, but adapting only $$Wq$$ or $$Wk$$ resulted in <mark style="color:yellow;">**significantly lower performance**</mark>.

This suggests that even with a <mark style="color:blue;">**low rank**</mark> (e.g., r=4), adapting multiple <mark style="color:blue;">**weight matrices**</mark> captures more useful information than adapting a <mark style="color:blue;">**single weight matrix**</mark> with a <mark style="color:blue;">**higher rank.**</mark>

<figure><img src="/files/zCsDYQh8u23pBYMBXp7N" alt=""><figcaption></figcaption></figure>

### <mark style="color:green;">**What is the optimal rank for the adaptation matrix ∆W**</mark>

The authors investigated the effect of the LoRA <mark style="color:blue;">**rank r**</mark> on downstream performance.

Surprisingly, they found that a <mark style="color:yellow;">**rank as low as r=1 was sufficient**</mark> for adapting both $$Wq$$ and $$Wv$$ on the datasets they tested, while adapting $$Wq$$ alone required a larger <mark style="color:blue;">**rank r**</mark>.

They further analysed the subspace similarity between the learned adaptation matrices with different ranks (r=8 and r=64) and found that the *<mark style="color:yellow;">**top singular vector directions overlapped significantly**</mark>*, while the other directions did not.&#x20;

This suggests that the additional directions learned with higher ranks might contain mostly random noise.

The authors conclude that the optimal <mark style="color:blue;">**adaptation matrix**</mark> ∆W can indeed have a very low intrinsic rank, although they note that this may not hold for every task or dataset.

### <mark style="color:green;">**Connection between ∆W and W**</mark>

To investigate the relationship between the <mark style="color:blue;">**adaptation matrix**</mark> $$∆W$$ and the <mark style="color:blue;">**pre-trained weight matrix**</mark> $$W$$, the authors projected $$W$$ onto the r-dimensional subspace of $$∆W$$ and compared the [<mark style="color:blue;">**Frobenius norms**</mark>](#user-content-fn-5)[^5].

They found that $$∆W$$ has a stronger correlation with $$W$$<mark style="color:yellow;">**compared to a random matrix**</mark>, indicating that $$∆W$$amplifies some features that are already present in $$W$$.&#x20;

However, instead of amplifying the top singular directions of  $$W$$, $$∆W$$ emphasises directions that are not as prominent in $$W$$.

The amplification factor is quite large (e.g., 21.5 for r=4 in the 48th layer of GPT-3).

This suggests that the low-rank adaptation matrix <mark style="color:yellow;">**amplifies important features for specific downstream tasks**</mark> that were *<mark style="color:yellow;">**learned but not emphasised in the general pre-training model**</mark>*.

### <mark style="color:green;">**Process for determining the optimal rank r for LoRA when fine-tuning**</mark>

1. Start with a low <mark style="color:blue;">**rank r**</mark> (e.g., r=1 or r=2) and fine-tune the model on the downstream task.
2. Gradually increase the <mark style="color:blue;">**rank r**</mark> (e.g., r=4, r=8) and compare the performance on a validation set.
3. If increasing the rank leads to significant improvements, continue increasing <mark style="color:blue;">**rank r**</mark> until the performance gains plateau or the computational cost becomes too high.
4. If the performance is already good with a low rank, try *<mark style="color:yellow;">**adapting additional weight matrices**</mark>* (e.g., $$Wq$$and $$Wv$$ together) with the same low rank.
5. Compare the performance and computational cost of different combinations of rank and adapted weight matrices to find the optimal configuration for the specific downstream task and resource constraints.

Keep in mind that the optimal rank may vary depending on the complexity of the downstream task, the size of the dataset, and the similarity between the pre-training and downstream domains.&#x20;

It's also important to consider the trade-off between performance and computational efficiency when choosing the rank.

### <mark style="color:purple;">Contents of A and B</mark>

The <mark style="color:blue;">**matrices**</mark> $$( A )$$ and $$( B )$$ contain learned parameters that are updated during the fine-tuning process. They are initialised randomly at the beginning of training:

* Matrix <mark style="color:blue;">**matrices**</mark> is initialised with a random Gaussian distribution.
* Matrix $$( B )$$is initialised with zeros.

During training, the values in $$( A )$$ and $$( B )$$  are updated based on the gradients computed during <mark style="color:blue;">**backpropagation**</mark>.&#x20;

These matrices learn to adapt the <mark style="color:blue;">**pre-trained weights**</mark> to the specific downstream task by capturing the important patterns and transformations needed for the adaptation.

The content of $$( A )$$ and $$( B )$$  is learned through the optimisation process and depends on the specific task and dataset being fine-tuned on.&#x20;

The learned values in these matrices represent the low-rank adaptation that modifies the pre-trained weights to better suit the downstream task.

<figure><img src="/files/O6v2jTzCSGBr2kpbfta5" alt=""><figcaption><p>Backpropagation is a fundamental algorithm used in training neural networks and other differentiable machine learning models. It is a method for efficiently calculating the gradients of the model's parameters with respect to the loss function. The goal of backpropagation is to update the model's parameters in a way that minimizes the difference between the predicted output and the desired output.</p></figcaption></figure>

### <mark style="color:purple;">Why LoRA is Better!</mark>

LoRA addresses the limitations of existing solutions by introducing a more efficient and effective approach to model adaptation:

<mark style="color:blue;">**No Inference Latency:**</mark> Unlike adapter layers, LoRA does not introduce additional depth to the model. The low-rank adaptation matrices can be <mark style="color:yellow;">**merged with the pre-trained weights after fine-tuning**</mark>, resulting in no extra inference latency compared to a fully fine-tuned model.

<mark style="color:blue;">**Compute and Memory Efficiency:**</mark> LoRA uses a low-rank representation to encode the task-specific parameter increments, significantly reducing the number of trainable parameters. This makes fine-tuning more compute- and memory-efficient, especially for large models.

<mark style="color:blue;">**Optimisation Stability:**</mark> Compared to prompt tuning, LoRA optimises the model parameters directly, which leads to more stable optimisation and monotonic performance improvements with increased model capacity.

<mark style="color:blue;">**Sequence Length Preservation:**</mark> LoRA does not require reserving a portion of the sequence length for adaptation, allowing the full sequence length to be used for downstream tasks, potentially leading to better performance.

<mark style="color:blue;">**Flexibility and Composability:**</mark> LoRA is agnostic to the training objective and can be easily integrated with existing models and architectures. It is also composable with other adaptation techniques, such as prefix tuning, offering further flexibility.

<mark style="color:blue;">**Enhanced Compatibility**</mark><mark style="color:blue;">:</mark> Works well alongside other fine-tuning techniques like adapters and prefix tuning.

## <mark style="color:purple;">Conclusion</mark> <a href="#id-3d50" id="id-3d50"></a>

LoRA (Low-Rank Adaptation) is a technique for fine-tuning large language models that takes advantage of the intrinsically low rank of the weight matrices.&#x20;

Instead of updating the entire weight matrix during fine-tuning, LoRA decomposes the weight update matrix into two smaller matrices (A and B) with a lower <mark style="color:blue;">**rank r**</mark>.&#x20;

This significantly reduces the number of trainable parameters and memory requirements while maintaining performance comparable to full fine-tuning.

### <mark style="color:purple;">Key Insights</mark>

#### <mark style="color:blue;">**Swappable LoRA Modules**</mark>

One of the most significant advantages of LoRA is the ability to swap different fine-tuned LoRA modules for various tasks on a single base model.&#x20;

This allows for a more flexible and efficient deployment of models, as you can easily switch between tasks without needing separate fully fine-tuned models for each task.

#### <mark style="color:blue;">**Inference Time Swapping**</mark>

The swappable nature of LoRA modules can be used even at inference time.&#x20;

This means that customers can choose which task they want the model to perform on-the-fly, without requiring multiple models to be running simultaneously. This is a powerful feature that sets LoRA apart from other adaptation methods.

#### <mark style="color:blue;">**Potential for Further Optimisation**</mark>

While LoRA is often applied to the attention weights (specifically query and value matrices) in transformer models, the technique could potentially be *<mark style="color:yellow;">**applied to other weight matrices in the model**</mark>*.&#x20;

Exploring the application of LoRA to different components of the model architecture could lead to further optimizations and improvements.

#### <mark style="color:blue;">**Balancing Rank and Performance**</mark>

The *<mark style="color:yellow;">**rank of the low-rank matrices (A and B) in LoRA is a crucial hyperparameter**</mark>* that determines the trade-off between model performance and efficiency. While lower ranks lead to greater memory savings and faster training, it's essential to find the right balance for each specific task. Experimenting with different rank values and evaluating the results on a validation set can help determine the optimal configuration.

#### <mark style="color:blue;">**Implications for Model Accessibility**</mark>

By significantly reducing the memory requirements and training costs, LoRA makes fine-tuning large language models more accessible to a wider range of researchers and practitioners.&#x20;

This could accelerate the development and deployment of specialized models for various tasks and domains.

<mark style="color:blue;">**Handling large datasets**</mark>

Fine-tuning with LoRA on large datasets may still require significant computational resources, even though the number of trainable parameters is reduced. Strategies such as data parallelism, model parallelism, or gradient accumulation can be employed to handle large datasets efficiently.

[^1]: Subspace refers to a smaller dimensional space within the larger parameter space of the model. Essentially, it means that the model's parameters can be <mark style="color:green;">**represented or projected onto a simpler and smaller set of dimensions**</mark>, allowing the model to perform effectively without needing the full complexity of the original parameter space.

[^2]: Gradient descent is an optimisation algorithm used to minimise the loss function by iteratively moving towards the direction of steepest descent as defined by the negative of the gradient. At each step, the model parameters are adjusted in the opposite direction of the gradient, scaled by a learning rate. This process continues until the model converges to a local minimum of the loss function. It's a fundamental technique used in training machine learning models. including neural networks.

[^3]: **WikiSQL** is a benchmark dataset for natural language to SQL translation. The task involves converting natural language queries into SQL queries that can be executed on a relational database.  <mark style="color:green;">**Example task:**</mark> Given the question "What is the population of France?" and a table of countries with their populations, the model needs to generate the correct SQL query to retrieve the population of France.

[^4]: **MultiNLI** is a dataset used for the natural language inference task. The goal is to determine the relationship between a pair of sentences: whether the second sentence (hypothesis) is entailed by, contradicts, or is neutral with respect to the first sentence (premise).

[^5]: The **Frobenius norm** of a matrix is a measure of its magnitude - calculated as the square root of the sum of the absolute squares of its elements.   This norm gives a sense of the overall size of the entries in the matrix and is used in optimisation and machine learning to compare magnitudes of different matrices.


# Practical Tips for Fine-tuning LMs Using LoRA (Low-Rank Adaptation)

This is a terrific article from the genius Sebastian Raschka, Phd.&#x20;

{% embed url="<https://magazine.sebastianraschka.com/p/practical-tips-for-finetuning-llms>" %}
An excellent article from [SEBASTIAN RASCHKA, PHD](https://substack.com/@rasbt)
{% endembed %}

The article provides valuable insights and lessons learned from the author's experiments with Low-rank Adaptation (LoRA), a widely used technique for efficiently training custom large language models (LLMs).&#x20;

The main takeaways include the consistency of LoRA training outcomes across multiple runs, the trade-off's of using QLoRA (quantized LoRA) for memory savings, and the minimal impact of optimizer choice on LLM fine-tuning.

The author also discusses the importance of applying LoRA across all layers, adjusting the LoRA rank and alpha value, and the feasibility of fine-tuning 7 billion parameter models on a single GPU.&#x20;

Additionally, the article addresses common questions related to LoRA, such as the significance of the dataset, the effectiveness of LoRA for domain adaptation, and strategies for avoiding overfitting.

The author compares LoRA to full fine-tuning and RLHF (Reinforcement Learning with Human Feedback), highlighting the memory efficiency and performance of LoRA.&#x20;

The article also explores the possibility of combining multiple sets of LoRA weights and discusses the concept of Layer-wise Optimal Rank Adaptation.

### <mark style="color:purple;">Here is a summary of the author's tips</mark>

<mark style="color:green;">**Consistency in LLM Training**</mark><mark style="color:green;">:</mark> Even though there's some randomness in training language models (LMs) and models on GPUs, the results are usually pretty consistent when you run the training multiple times.

<mark style="color:green;">**QLoRA for Memory Efficiency**</mark><mark style="color:green;">:</mark> QLoRA is a good choice when you don't have a lot of GPU memory. It can save about a third of the memory but makes training about 39% slower. It's a good option if memory is your biggest problem.

<mark style="color:green;">**Optimiser Choice in Fine-Tuning**</mark><mark style="color:green;">:</mark> It doesn't make a big difference which optimiser you choose (AdamW, SGD with scheduler, AdamW with scheduler). SGD by itself isn't as good, but the others are all pretty similar.

<mark style="color:green;">**Adam Optimizer and Memory Usage**</mark><mark style="color:green;">:</mark> The Adam optimizer uses more memory because it has two extra numbers for each number in the model. But for LMs, this doesn't make the memory usage a lot higher because most of the memory is used for big matrix calculations, not for storing the extra numbers.

<mark style="color:green;">**Multi-Epoch Training and Static Datasets**</mark><mark style="color:green;">:</mark> Training on the same dataset multiple times (multi-epoch training) might not help and could even make the model worse, probably because it starts to overfit the data.

<mark style="color:green;">**Application of LoRA**</mark><mark style="color:green;">:</mark> To make the model work its best, use LoRA on all the layers, not just the Key and Value matrices.

<mark style="color:green;">**Adjusting LoRA Parameters**</mark><mark style="color:green;">:</mark> It's important to choose the right LoRA rank and alpha value. A good rule of thumb is to make alpha twice as big as the rank.

<mark style="color:green;">**Fine-tuning 7 Billion Parameter Models**</mark><mark style="color:green;">:</mark> You can finetune these big models in just a few hours on a single GPU with 14 GB of memory. But it's hard to make an LLM do well on all benchmark tasks with just one dataset. You might need to use different datasets or tools.

### <mark style="color:purple;">**Expanding LoRA to More Layers**</mark>

The article talks about experiments where LoRA was first used only on Key and Value weight matrices in transformer layers.

Using it on Query weight matrices, projection layers, and other linear layers too makes the number of trainable parameters much bigger (from about 4.2 million to over 20 million for a model with 7 billion parameters).

This uses more memory (16.62 GB instead of 14.18 GB) but can make the model perform noticeably better.&#x20;

The author says they only tried two settings (LoRA for just the query and value matrices, and LoRA for all layers) and suggests that future experiments should look at other combinations, like what happens if you use LoRA for projection layers.

<mark style="color:green;">**Balancing LoRA Hyperparameters - Rank (R) and Alpha (α)**</mark>

The article explains the importance of the scaling coefficient in LoRA, which uses the rank parameter (r) and another hyperparameter α (alpha).&#x20;

The formula for scaling is α / r, and the LoRA weights' influence gets bigger with this scaling factor. The author tried different rank values and found that making α twice as big as r usually gives the best results. This was especially clear when r was set to 256, where the best α was found to be 512.

<mark style="color:green;">**Training 7 Billion Parameter Models on a Single GPU**</mark>

One of the big benefits of LoRA, as the article points out, is that it lets you fine-tune big models (like a model with 7 billion parameters) on just one GPU.&#x20;

Using QLoRA with the best settings (r=256 and α=512) and an AdamW optimizer, you can fine-tune a model this big in about 3 hours on an A100 GPU, even with a big training dataset (like the Alpaca dataset with 50,000 examples).


# QLORA: Efficient Finetuning of Quantized LLMs

This paper introduces QLORA, a parameter efficient fine tuning approach customising large language models (LLMs) with significantly reduced memory requirements.&#x20;

QLORA combines 4-bit quantization of the pretrained model with Low Rank Adapters (LoRA) to enable finetuning of a 65B parameter model on a single 48GB GPU, without sacrificing performance compared to full 16-bit finetuning.

{% embed url="<https://arxiv.org/abs/2305.14314>" %}

### <mark style="color:purple;">Key innovations of QLORA include</mark>

<mark style="color:blue;">**4-bit NormalFloat (NF4):**</mark> An information-theoretically optimal data type for normally distributed weights.

<mark style="color:blue;">**Double Quantization:**</mark> Quantizing the quantization constants to reduce memory footprint further.

<mark style="color:blue;">**Paged Optimizers:**</mark> Managing memory spikes during training using NVIDIA unified memory.

The authors use QLORA to finetune over 1,000 models, demonstrating state-of-the-art results with their Guanaco model family.&#x20;

Guanaco reaches 99.3% of ChatGPT's performance on the Vicuna benchmark while being trainable on a single GPU.

<mark style="color:green;">**The extensive analysis reveals several key findings**</mark>

1. <mark style="color:yellow;">Data quality is more important than dataset size</mark> for instruction finetuning and chatbot performance.
2. Strong performance on the MMLU benchmark does not necessarily imply strong chatbot performance, highlighting the <mark style="color:yellow;">importance of task-specific datasets</mark>.
3. GPT-4 <mark style="color:yellow;">evaluations largely agree with human evaluations in ranking chatbot performance</mark>, offering a cheaper alternative to human annotation, albeit with some uncertainties.

The authors release their codebase, CUDA kernels, and integrate their methods into the Hugging Face transformers library, making QLORA accessible to the community. They also release 32 finetuned models across various sizes and instruction datasets.

In summary, the QLORA paper introduces a groundbreaking approach to efficiently finetune large language models, democratising access to LLM finetuning and enabling in-depth analysis of instruction finetuning and chatbot performance at unprecedented scales.&#x20;

The open-source release of the code and models further contributes to the advancement of the field.

### <mark style="color:purple;">Background</mark>

To provide more context on quantization and its mathematical foundations, let's dive deeper into the background and explain dequantization and the potential risks involved.

Quantization is a technique used to reduce the precision of numerical representations, typically by mapping a larger set of values to a smaller set.&#x20;

In the context of deep learning, quantization is often applied to <mark style="color:blue;">**model weights**</mark> and <mark style="color:blue;">**activations**</mark>, converting them from higher-precision data types (e.g., 32-bit floating-point) to lower-precision data types (e.g., 8-bit integers).

This reduces memory consumption and can accelerate computations, especially on hardware optimized for lower-precision arithmetic.

<figure><img src="/files/T7qx7kpXOgUBiYKM9Mzr" alt=""><figcaption></figcaption></figure>

### <mark style="color:green;">Block-wise k-bit Quantization</mark>

The quantization process involves scaling the input values to fit within the range of the target data type.&#x20;

For example, when quantizing a 32-bit floating-point tensor to an 8-bit integer tensor with a range of \[-127, 127], the quantization formula is:

$$
XInt8 = round(127 / absmax(XFP32) \* XFP32)
$$

Here, $$absmax(XFP32)$$ represents the <mark style="color:blue;">**absolute maximum value**</mark> in the input tensor.

The <mark style="color:yellow;">**scaling factor,**</mark> $$127 / absmax(XFP32)$$, is called the <mark style="color:blue;">**quantization constan**</mark>t or quantization scale, denoted as c.

To mitigate the impact of outliers on the quantization process, block-wise quantization is employed.

&#x20;The <mark style="color:yellow;">**input tensor is divided into smaller blocks**</mark>, and each block is quantized independently with its own quantization constant.&#x20;

This ensures better utilization of the available quantization bins.

#### <mark style="color:green;">**Dequantization**</mark>

Dequantization is the inverse process of quantization, where the quantized values are mapped back to their original data type. The dequantization formula for the example above is:

$$
XFP32 = XInt8 / c
$$

Here, $$c$$ is the <mark style="color:yellow;">**quantization constant**</mark> used during the quantization step.

Risks and Considerations:

<mark style="color:blue;">**Information Loss:**</mark> Quantization inherently leads to a loss of information due to the reduced precision. This can affect the model's accuracy and performance, especially if the quantization is too aggressive.

<mark style="color:blue;">**Quantization Noise:**</mark> The quantization process introduces noise into the model, as the original values are approximated by the quantized values. This noise can accumulate across layers and impact the model's behavior.

<mark style="color:blue;">**Outliers and Range:**</mark> Outliers in the input tensor can significantly affect the quantization process, leading to poor utilization of the available quantization bins. Block-wise quantization helps mitigate this issue, but it's still important to consider the range of values in the tensor.

<mark style="color:blue;">**Hardware Compatibility:**</mark> While quantization can lead to memory savings and computational speedups, the target hardware must support the specific quantized data types and operations. Not all hardware platforms have efficient support for low-precision arithmetic.

<mark style="color:blue;">**Quantization-Aware Training:**</mark> To achieve optimal performance with quantized models, quantization-aware training techniques can be employed. These techniques simulate the quantization process during training, allowing the model to adapt to the quantization noise and minimize its impact on accuracy.

Despite these risks, quantization remains a powerful technique for reducing the memory footprint and computational requirements of deep learning models.&#x20;

By carefully considering the trade-offs and employing appropriate quantization strategies, such as block-wise quantization and quantization-aware training, the impact of quantization on model performance can be minimized while realizing significant efficiency gains.

<figure><img src="/files/NQOY6qBI5rOAJKBFMzpJ" alt=""><figcaption></figcaption></figure>

### <mark style="color:purple;">Different components of QLORA</mark>

<mark style="color:green;">**4-bit NormalFloat Quantization**</mark>

The authors observe that pretrained neural network weights usually follow a zero-cantered normal distribution with a standard deviation σ.&#x20;

This means that the *<mark style="color:yellow;">**weights are symmetrically distributed around zero**</mark>*, and the spread of the distribution is determined by the standard deviation.

To optimises the quantization process for such normally distributed weights, they introduce the 4-bit NormalFloat (NF4) quantization.&#x20;

The idea is to create a quantization scheme that is information-theoretically optimal for zero-mean normal distributions.&#x20;

The process involves:&#x20;

<mark style="color:blue;">a.</mark> Estimating the $$2^k + 1$$quantiles of a standard normal distribution $$N(0,1)$$to obtain a k-bit quantile quantization data type

<mark style="color:blue;">b.</mark> Normalizing the data type values into the range $$\[-1, 1]$$

<mark style="color:blue;">c.</mark> Quantizing the input weight tensor by normalizing it into the range \[-1, 1] using absolute maximum rescaling.

The equation $$qi = 1/2 \* (QX(i/(2^k+1)) + QX((i+1)/(2^k+1)))$$estimates the quantile values $$qi$$ for the data type, where $$QX(·)$$is the quantile function of the standard normal distribution.

To ensure an exact representation of zero, they create an asymmetric data type by estimating the quantiles separately for the negative and positive parts and then unifying the sets while removing one of the duplicate zeros.

<mark style="color:green;">**Double Quantization**</mark>

Double Quantization (DQ) is introduced to reduce the memory footprint of the quantization constants. It involves *<mark style="color:yellow;">**quantizing the quantization constants themselves**</mark>*.

The process works as follows:&#x20;

<mark style="color:blue;">**a.**</mark> The quantization constants $$cFP32$$from the first quantization are treated as inputs to a second quantization.&#x20;

<mark style="color:blue;">**b.**</mark> The second quantization yields the quantized quantization constants $$cFP8$$ and the second level of quantization constants $$cFP32$$.&#x20;

<mark style="color:blue;">**c.**</mark> 8-bit Floats with a blocksize of 256 are used for the second quantization to avoid performance degradation.&#x20;

<mark style="color:blue;">**d.**</mark> Since $$cFP32$$ values are positive, the mean is subtracted from $$c2$$before quantization to centre the values around zero and enable symmetric quantization.

This double quantization reduces the memory footprint per parameter from 0.5 bits to 0.127 bits, achieving a reduction of 0.373 bits per parameter.

<mark style="color:green;">**QLORA**</mark>

QLORA combines the 4-bit NormalFloat quantization, Double Quantization, and Low-Rank Adapters (LoRA) to achieve efficient 4-bit quantization.

For a single linear layer in the quantized base model with a single LoRA adapter, QLORA is defined as:

$$
YBF16 = XBF16 \* doubleDequant(cFP32\_1, ck-bit\_2, WNF4) + XBF16 \* LBF16
$$

where <mark style="color:blue;">**doubleDequant(·)**</mark> is the double dequantization process:&#x20;

$$
doubleDequant(cFP32\_1, ck-bit\_2, Wk-bit) =
$$

$$
dequant(dequant(cFP32\_1, ck-bit\_2), W4bit) = WBF16
$$

QLORA uses NF4 for the weights $$(W)$$ and FP8 for the quantization constants $$(c2)$$.&#x20;

The blocksize is set to 64 for W for higher precision and 256 for c2 to conserve memory.

During the backward pass, only the gradients with respect to the LoRA adapter weights $$(∂E/∂Li)$$are computed, not for the 4-bit weights $$(∂E/∂W)$$.&#x20;

However, computing $$∂E/∂Li$$ involves calculating $$∂X/∂W\* e^{2 pi i \xi x}$$, which requires dequantizing the storage WNF4 to the computation data type WBF16.

In summary, QLORA uses 4-bit NormalFloat as the storage data type and 16-bit BrainFloat as the computation data type.&#x20;

The storage data type is dequantized to the computation data type for the forward and backward passes, but gradients are only computed for the LoRA parameters in 16-bit precision.

### <mark style="color:purple;">QLORA vs. Standard Finetuning</mark>

To compare QLORA with standard finetuning, the authors conduct experiments on various architectures (encoder, encoder-decoder, and decoder-only) and model sizes (up to 3B parameters).&#x20;

### <mark style="color:green;">**Best Practices**</mark>

<mark style="color:blue;">**LoRA Adapters:**</mark> The authors find that applying LoRA to all linear transformer block layers is crucial to match the performance of full finetuning. The number of LoRA adapters used is the most critical hyperparameter.

<mark style="color:blue;">**Hyperparameter Tuning:**</mark> Default hyperparameters for fully finetuned baselines are often undertuned. The authors perform a hyperparameter search over <mark style="color:yellow;">learning rates</mark> (1e-6 to 5e-5) and <mark style="color:yellow;">batch sizes</mark> (8 to 128) to establish robust baselines.

### <mark style="color:green;">**Comparison**</mark>

#### <mark style="color:blue;">**4-bit NormalFloat (NF4) vs. 4-bit Floating Point (FP4)**</mark>

NF4 significantly improves performance over FP4 and Int4 data types. Double quantization reduces the memory footprint without degrading performance.

#### <mark style="color:blue;">**QLORA vs. 16-bit Full Finetuning and 16-bit LoRA**</mark>

The authors find that 4-bit QLORA with the NF4 data type matches the performance of both 16-bit full finetuning and 16-bit LoRA finetuning on academic benchmarks. This holds true for various model sizes (125M to 65B parameters) and datasets (GLUE, Super-Natural Instructions, Alpaca, and FLAN v2).

#### <mark style="color:blue;">Performance-Precision Trade-off</mark>

In line with previous work on quantization, the authors observe that with a given finetuning and inference resource budget, it is beneficial to increase the number of parameters in the base model while decreasing their precision. This highlights the importance of efficiency benefits from QLORA.

<mark style="color:green;">**Key Findings**</mark>

1. QLORA with NF4 <mark style="color:yellow;">replicates both 16-bit full finetuning and 16-bit LoRA finetuning performance</mark>.
2. <mark style="color:yellow;">NF4 is superior to FP4</mark> in terms of quantization precision.
3. <mark style="color:yellow;">Double quantization does not degrade performance</mark>.

The authors' results consistently show that 4-bit QLORA with the NF4 data type matches the performance of 16-bit methods while offering significant memory savings.

This allows for the exploration of instruction tuning at scales that would be impossible with full 16-bit finetuning on academic research hardware.

### <mark style="color:purple;">Limitations</mark>

<mark style="color:green;">**Lack of comparison with full 16-bit finetuning at larger scales**</mark>

While the authors provide evidence that QLORA can replicate 16-bit full finetuning performance with a 4-bit base model and Low-rank Adapters (LoRA), they did not establish this at the 33B and 65B scales due to the immense resource costs involved.

#### <mark style="color:green;">Limited evaluation on instruction finetuning models</mark>

The authors evaluated QLORA on MMLU, the Vicuna benchmark, and the OA benchmark but did not evaluate on other benchmarks such as BigBench, RAFT, and HELM. It is not ensured that their evaluations generalize to these other benchmarks.

<details>

<summary><mark style="color:green;">Benchmarks?</mark></summary>

The performance of models against these benchmarks is measured using various methods and metrics, depending on the specific focus of each benchmark. Here is an overview of how performance is typically measured for each:

<mark style="color:blue;">**MMLU (Massive Multitask Language Understanding)**</mark>

* **Measurement Method:** Multiple-choice questions.
* **Metrics:** Accuracy is the primary metric, calculated as the percentage of correct answers out of the total questions.
* **Details:** Performance is evaluated across 57 tasks from different domains, and the overall accuracy provides a comprehensive measure of the model's general knowledge and understanding.

<mark style="color:blue;">**Vicuna Benchmark**</mark>

* **Measurement Method:** Evaluation of conversational tasks and scenarios.
* **Metrics:** Human evaluation scores, BLEU (Bilingual Evaluation Understudy), ROUGE (Recall-Oriented Understudy for Gisting Evaluation), and other dialogue-specific metrics.
* **Details:** Human judges often rate the quality of responses based on coherence, relevance, informativeness, and fluency. Automated metrics may also be used to compare generated text against reference responses.

<mark style="color:blue;">**OA (OpenAI) Benchmark**</mark>

* **Measurement Method:** A variety of tasks designed to test different capabilities.
* **Metrics:** Task-specific metrics such as accuracy, F1 score, precision, recall, and others depending on the nature of each task.
* **Details:** The benchmark includes diverse tasks, and the performance is measured using the appropriate metric for each task to provide a detailed view of the model's strengths and weaknesses.

<mark style="color:blue;">**BigBench (Beyond the Imitation Game Benchmark):**</mark>

* **Measurement Method:** A wide range of tasks developed by the research community.
* **Metrics:** Varies by task; common metrics include accuracy, F1 score, and others relevant to the specific task.
* **Details:** The benchmark covers reasoning, commonsense understanding, and other advanced skills. Performance is evaluated task by task, and an aggregate score may be used to summarize overall performance.

<mark style="color:blue;">**RAFT (Realistic Adversarial Functionality Test)**</mark>

* **Measurement Method:** Adversarial tasks designed to expose model weaknesses.
* **Metrics:** Accuracy, robustness metrics, error rates, and other task-specific metrics.
* **Details:** Models are tested with challenging and tricky inputs to assess their robustness and reliability. Performance is measured by how well the model can handle these difficult scenarios.

<mark style="color:blue;">**HELM (Holistic Evaluation of Language Models)**</mark>

* **Measurement Method:** A comprehensive set of evaluations across various dimensions.
* **Metrics:** Accuracy, fairness metrics, robustness metrics, efficiency (e.g., speed, computational resources), and others.
* **Details:** The benchmark aims to provide a holistic view of performance, considering multiple aspects beyond just accuracy. Metrics are chosen to reflect the model's performance in terms of fairness, robustness, and efficiency.

In summary, each benchmark employs specific methods and metrics tailored to its focus area, providing a nuanced and detailed assessment of language model performance across different tasks and dimensions.

</details>

#### <mark style="color:green;">**Dependency on similarity between finetuning data and benchmark data**</mark>

The performance of the benchmarks likely depends on how similar the finetuning data is to the benchmark dataset. This highlights the need for better benchmarks and evaluation methods, as well as careful consideration of what is being evaluated in the first place.

#### <mark style="color:green;">Limited responsible AI evaluation</mark>

&#x20;While the authors evaluate the likelihood of Guanaco-65B to generate a socially biased sequence of tokens compared to other models, it is unclear if Guanaco performs well when assessed on other types of biases.


# Bits and Bytes

Tim Dettmers (PhD candidate, University of Washington) presents "8-bit Methods for Efficient Deep Learning" in this Cohere For AI Technical Talk.

Language models are effective tools for many tasks but are difficult to train and inference due to their size.&#x20;

Moving from 32-bit models to 16-bit models resulted in considerable efficiency gains that made training and inference of large models easier.&#x20;

Can we train and inference in 8-bit to make further gains?&#x20;

In this talk, Tim will show that <mark style="color:yellow;">8-bit inference and training can be used without degrading performance while improving efficiency.</mark>&#x20;

To make 8-bit methods work, it is essential to understand how quantization precision affects model performance and training stability as we scale the model size.&#x20;

He will talk about how these factors change with scale and how we need to adjust 8-bit methods to make them work.&#x20;

In particular, he will speak about 8-bit optimizers for training and Int8 inference for large language models with up to 175B parameters. These methods make training and inference more efficient and make large models more accessible to researchers.

{% embed url="<https://www.youtube.com/watch?v=jyOqtw4ry2w>" %}
Bits and Bytes
{% endembed %}

### <mark style="color:purple;">Summary of Transcript</mark>

<mark style="color:blue;">Quantization</mark> is the process of mapping a large set of input values to a smaller set of discrete values, similar to histogram binning.&#x20;

In <mark style="color:blue;">linear quantization</mark> (integer quantization), the input range is divided into equal-sized bins, and each value within a bin is mapped to the bin's middle value.

<mark style="color:blue;">Non-linear quantization</mark> allows for varying bin widths, providing higher precision in certain regions of the input range.

Tim introduces a dynamic exponent datatype that efficiently represents a wide range of values by allocating bits dynamically between the exponent and fraction parts.  This datatype is particularly useful for representing extreme values (very large or very small) with high precision.

The talks about 8-bit optimizers, which reduce the memory footprint of training by quantizing the optimizer states (e.g., Adam optimizer's momentum and velocity buffers) to 8 bits.&#x20;

However, outliers in the optimizer states can lead to significant quantization errors.  To mitigate this, Tim proposes chunking the optimizer states into blocks and treating each block independently, isolating the impact of outliers. This method achieves performance similar to 32-bit optimizers while reducing memory usage.

Next, Tim discusses LLM.int8, a method for efficient inference of large language models using 8-bit quantization.&#x20;

Outliers in activations can cause significant performance degradation in 8-bit quantized models.&#x20;

LLM.int8 addresses this by identifying outlier-prone columns in the activations and processing them in 16-bit precision while keeping the rest of the activations in 8-bit precision.&#x20;

This approach maintains the performance of 16-bit models while reducing memory usage, making large models like OPT-175B and LLaMA-65B accessible on consumer hardware.

Finally, Tim presents his recent work on optimal quantization for inference, comparing the performance of models with varying bit-widths and parameter counts.&#x20;

Through extensive experiments, he finds that <mark style="color:yellow;">4-bit quantization provides the best balance between model size and performance</mark>.&#x20;

Models with 4-bit weights and 16-bit activations consistently outperform models with higher bit-widths and fewer parameters. Tim also explores the impact of block size and datatype on quantization performance, showing that smaller block sizes (e.g., 64) and floating-point or quantile-based datatypes yield the best results.

### <mark style="color:purple;">Tips and Tricks</mark>

1. When quantizing models, consider the distribution of your data and <mark style="color:yellow;">choose appropriate bin widths to minimize quantization errors.</mark>
2. Use <mark style="color:yellow;">block-wise quantization</mark> to isolate the impact of outliers and improve quantization stability.
3. For inference, <mark style="color:yellow;">4-bit quantization provides the best balance between model size and performance.</mark> Use 4-bit weights and 16-bit activations for optimal results.
4. <mark style="color:yellow;">Experiment with different block sizes and datatypes</mark> to further optimize quantization performance. Smaller block sizes and floating-point or quantile-based datatypes tend to yield better results.
5. Be aware of the <mark style="color:yellow;">trade-offs between quantization precision and model size</mark>. Lower bit-widths may require more parameters to achieve the same performance as higher bit-widths.

In conclusion, quantization techniques are powerful tools for reducing the memory footprint and computational cost of deep learning models.&#x20;

By carefully choosing quantization schemes, datatypes, and block sizes, you can achieve significant memory savings while maintaining high performance. As demonstrated by Tim Detmers' work, these techniques are particularly valuable for making large language models more accessible and efficient.


# The Magic behind Qlora

### <mark style="color:purple;">Introduction</mark>

Language Models (LMs) have revolutionised the field of natural language processing, enabling breakthroughs in tasks such as language understanding, generation, and translation.&#x20;

These models, built on the transformer architecture, consist of layers with multi-head self-attention mechanisms and feed-forward neural networks. Each component has associated weight matrices that are crucial for the model's functioning and can have millions or billions of parameters in large-scale models.

Traditional fine-tuning approaches involve updating the entire weight matrix of a model, which can be computationally expensive and memory-intensive.&#x20;

To address this challenge, methods like Lora (Low-Rank Adaptation) and Qlora (Quantized Lora) have emerged, focusing on updating a smaller, decomposed gradient matrix. This shift allows for more efficient training by updating specific, newly added, or adapted components instead of retraining the entire model.

{% embed url="<https://www.youtube.com/watch?v=DZoICV92VGE>" %}
The Magic behind Qlora
{% endembed %}

### <mark style="color:purple;">Low-Rank Adapters and Matrix Decomposition</mark>&#x20;

At the core of Lora and Qlora is the concept of <mark style="color:blue;">embedded low-rank adapters</mark>, which rely on matrix decomposition.&#x20;

These adapters are small neural network layers inserted between existing layers of a pre-trained model, designed to capture the essential transformations needed for a new task while keeping the complexity low.  The term "low rank" refers to the fact that these adapters have fewer parameters compared to the main layers of the model.

<mark style="color:blue;">Matrix decomposition</mark> involves representing a high-rank matrix (the weight or gradient matrix) with a lower-rank matrix (the adapter).&#x20;

This process decomposes a complex, high-dimensional matrix into a simpler form that still captures the essential information. By focusing on updating the low-rank adapters instead of the entire weight matrix, Lora and Qlora enable more efficient fine-tuning of LLMs.

### <mark style="color:purple;">Quantization and Information Loss Mitigation</mark>&#x20;

Qlora takes the concept of low-rank adaptation further by introducing quantization techniques, including the use of a new data type called the 4-bit normal float.&#x20;

Quantization involves mapping continuous or high-precision values to a smaller set of discrete values, reducing memory usage at the cost of some information loss.

However, Qlora employs an innovative approach to mitigate this information loss.&#x20;

Unlike traditional quantization methods that use a single constant, Qlora computes separate quantization constants for each block of weights. This approach ensures a more accurate and nuanced quantization, minimizing the loss of information and handling the distribution of weights more effectively.

The <mark style="color:blue;">quantization constant</mark> represents the scaling factor used in the quantization process, where the maximum value in a vector is scaled to fit within the quantized range.&#x20;

This constant is crucial for both the quantization and subsequent dequantization processes, ensuring that the original data can be closely approximated after being compressed.

### <mark style="color:purple;">Double Quantization and Efficiency</mark>

Qlora introduces the concept of <mark style="color:blue;">double quantization</mark>, which involves quantizing not just the model parameters but also the quantization constants themselves.  This approach further reduces the memory footprint, allowing for more efficient storage and processing of large models.

Compared to other parameter-efficient methods like prefix tuning and adapters, Qlora and Lora have shown superior performance, achieving comparable or better results with significantly fewer parameters. This effectiveness highlights the potential of these methods in efficiently fine-tuning LLMs.

### <mark style="color:purple;">Technical Details and Training</mark>&#x20;

The implementation of Qlora involves several technical considerations to optimize performance and stability.&#x20;

Model preprocessing steps, such as <mark style="color:yellow;">**upcasting layer norms to float32**</mark>, are employed to ensure more stable training. Memory management techniques, including the use of page optimisers and NVIDIA's unified memory, help address memory spikes that can occur when processing large mini-batches or long sequence lengths.

Gradient checkpointing is another crucial technique used in Qlora to balance memory usage and computational speed during backpropagation. By storing necessary activations for computing gradients without keeping all activations in memory, gradient checkpointing reduces the memory burden while allowing for efficient gradient computation.

Despite its advancements, Qlora maintains compatibility with standard optimization techniques like AdamW, ensuring seamless integration with existing training pipelines and optimization strategies.

### <mark style="color:purple;">Understanding Gradient Updates and Scalability</mark>&#x20;

To effectively apply Qlora, it is essential for researchers and developers to understand how gradient updates work in the context of low-rank adaptation.&#x20;

Visualizing these updates as modifications to a lower-rank matrix rather than the full weight matrix offers a more intuitive grasp of the process.

The scalability and reduced memory usage of Qlora make it highly applicable in industry settings. It allows for training multiple models for different tasks using a single base model, which is especially beneficial for tasks requiring frequent model updates, such as in e-commerce or recommendation systems.

Qlora's implementation is relatively straightforward, especially with tools like the Hugging Face library abstracting much of the complexity.&#x20;

This simplicity enhances the accessibility of the method to a wider range of developers and researchers. Moreover, Qlora is not limited to just the largest models; it can be applied to a range of model sizes, offering efficient fine-tuning capabilities across various architectures and scales.

### <mark style="color:purple;">Hyperparameter Experimentation and Future Directions</mark>&#x20;

The Qlora framework allows for experimentation with various hyperparameters, such as the <mark style="color:yellow;">number of Lora adapters, dropout rates, and layer-specific adaptations.</mark>&#x20;

This flexibility enables practitioners to fine-tune models to their specific needs and constraints. The analysis suggests that the <mark style="color:yellow;">number of Lora adapters used could become a significant hyperparameter</mark> in future implementations, with the goal of finding the optimal number that fits within specific GPU memory constraints while maximizing model performance.

The <mark style="color:yellow;">rank of the low-rank matrices is another critical hyperparameter in Qlora</mark>. A lower rank means fewer trainable parameters, which can greatly reduce the computational burden. However, finding the optimal rank involves balancing efficiency and maintaining the performance of the model.

The integration of Lora and Qlora weights into the LLM is a significant design choice, involving strategically placing these weights in different parts of the model to optimise performance while maintaining efficiency.&#x20;

This adaptability to various model architectures and tasks highlights the versatility of the approach.

### <mark style="color:purple;">Conclusion</mark>&#x20;

Qlora represents a significant advancement in the efficient fine-tuning of language models.&#x20;

By leveraging low-rank adaptation, quantization techniques, and innovative memory management strategies, Qlora enables the training of LMs with reduced computational and memory requirements.&#x20;

The method's scalability, compatibility with existing optimization techniques, and strong performance compared to other parameter-efficient methods make it a promising approach for both research and industry applications.

As the field of natural language processing continues to evolve, methods like Qlora will play a crucial role in making the development and deployment of large-scale models more accessible and efficient.&#x20;

Further research into hyperparameter optimisation, architectural integration, and applications across various tasks and domains will help unlock the full potential of these techniques.


# Practical Guide to LoRA: Tips and Tricks for Effective Model Adaptation

A range of practical tips and questions around using Lora

Fine-tuning large language models for specific tasks can significantly improve their performance.&#x20;

The Low Rank Adaptation (LoRA) technique offers an efficient pathway to achieve this without the extensive computational cost typically associated with full model fine-tuning.&#x20;

This guide outlines technical strategies and insights for effectively employing LoRA in model adaptation.

### <mark style="color:purple;">Key Strategies for LoRA Adaptation</mark>

#### <mark style="color:green;">Targeted Adaptation Focus</mark>

* Prioritise adapting the query and value weight matrices, either independently or alongside other weights, for enhanced performance.
* **Layer Selection**: Initial studies suggest that <mark style="color:yellow;">focusing on query and value matrices yields the best outcomes.</mark> You should consider various layer combinations to identify the most effective strategy.

#### <mark style="color:green;">Rank Selection and Efficiency</mark>

* **Exploring Low Ranks**: Even a rank of 1, turning matrices A and B into vectors, can be effective, suggesting that minimal parameter increases can still yield significant performance benefits.
* **Subspace Similarity Insights**: The top singular vector of a lower rank shows significant overlap with higher ranks, indicating that even low ranks capture critical higher-dimensional space information.

#### <mark style="color:green;">Domain-Specific Adaptation</mark>

* **Knowledge Absorption**: Leverage LoRA for domain-specific pretraining, especially when memory efficiency is crucial.
* **Task Diversity Consideration**: The diversity of tasks might necessitate larger ranks. This requires further investigation to establish a robust heuristic for rank selection based on the LLM and dataset in question.

#### <mark style="color:green;">Mitigating Overfitting</mark>

* **Rank and Overfitting**: Higher ranks may increase the risk of overfitting due to the expansion of trainable parameters.
* **Strategies for Mitigation**: Address overfitting by adjusting the rank, enlarging the dataset, modifying weight decay rates, or altering dropout rates specifically for LoRA layers.

#### <mark style="color:green;">Optimization Techniques</mark>

* **Sophia Optimizer**: Consider exploring the Sophia optimizer, known for its efficiency and performance benefits over traditional methods like Adam, especially for LLMs.

### <mark style="color:purple;">Practical Considerations</mark>

#### <mark style="color:green;">Memory Management</mark>

* **Influencing Factors**: Precision, quantization settings, model size, batch size, the number of trainable LoRA parameters, and dataset size all affect memory usage.
* **Sequence Length Optimization**: Shorter training sequences can lead to substantial memory savings, a vital consideration for managing computational resources.

#### <mark style="color:green;">Advanced Adaptation Techniques</mark>

* **Merging LoRA Weights**: It's feasible to combine multiple sets of LoRA weights for various applications, supported by tools like `merge_lora.py`.
* **Layer-Wise Rank Adaptation**: Analogous to selecting different learning rates for various layers, choosing distinct LoRA ranks for different layers adds a layer of customization but also complexity to the fine-tuning process.

### <mark style="color:purple;">Additional Insights</mark>

* **Efficient Model Adaptations**: Besides LoRA, adding adapter layers or optimizing input layer activations presents strategies for efficient model adaptation, each with its limitations, such as increased inference latency or optimization challenges.
* **Task Flexibility and Training Efficiency**: LoRA's design not only facilitates task flexibility, allowing a single pre-trained model to be adapted for multiple tasks, but also enhances training efficiency and inference performance without introducing additional latency.

LoRA emerges as a powerful tool for fine-tuning LLMs, offering a balance between computational efficiency and task-specific performance.&#x20;

By strategically selecting weights for adaptation, optimizing ranks, and managing computational resources, practitioners can leverage LoRA to enhance LLMs for a wide range of applications.


# The quantization constant

The <mark style="color:blue;">**quantization constant**</mark> in the context of QLORA (or any neural network quantization process) plays a crucial role in effectively reducing the model size while maintaining its performance.

### <mark style="color:purple;">**Role of the Quantization Constant**</mark>

* <mark style="color:green;">**Scaling Factor**</mark><mark style="color:green;">:</mark> The quantization constant acts as a scaling factor. During quantization, the continuous or high-precision values (like weights in a neural network) are converted into a more compact format. The quantization constant determines how these values are scaled down to fit within the limited range of the quantized format (e.g., 4-bit integers).
* <mark style="color:green;">**Maximising Use of Range**</mark><mark style="color:green;">:</mark> By scaling the maximum value in a vector to align with the quantized range, the quantization constant ensures that the available range is used optimally. This helps in maintaining the relative differences in the values, which is crucial for preserving the behavior of the neural network.

### <mark style="color:purple;">**Advantages of Separate Constants for Each Weight Block in QLORA**</mark>

* <mark style="color:green;">**Accuracy and Nuance**</mark><mark style="color:green;">:</mark> Different blocks of weights in a neural network may have different distributions. Using a single quantization constant for the entire network might not be optimal for all weight blocks. By computing separate quantization constants for each block, QLORA can tailor the quantization process to the specific distribution of each block, leading to more accurate quantization.
* <mark style="color:green;">**Minimising Information Loss**</mark><mark style="color:green;">:</mark> This tailored approach helps in minimizing the loss of information that typically occurs during quantization. Since each block is quantized according to its own characteristics, crucial details are less likely to be lost in the compression process.

### <mark style="color:purple;">**Importance in Dequantization**</mark>

* <mark style="color:green;">**Recovering Original Data**</mark><mark style="color:green;">:</mark> In the dequantization process, the quantized values are scaled back to their original range. The quantization constant is essential here to accurately reconstruct the original values.
* <mark style="color:green;">**Approximating Original Data**</mark><mark style="color:green;">:</mark> While exact recovery of the original data is often not possible due to the lossy nature of quantization, the use of the quantization constant allows for a close approximation, which is vital for maintaining the performance of the neural network.

In essence, the quantization constant in QLORA's approach is pivotal for efficiently compressing the neural network without significantly compromising its effectiveness.&#x20;

By customizing this constant for each weight block, QLORA enhances the precision of the quantization process, thereby achieving a balance between model compactness and performance retention.


# QLORA: Efficient Finetuning of Quantized Language Models

QLORA: Efficient Finetuning of Quantized Language Models

### <mark style="color:purple;">Introduction</mark>&#x20;

The process of fine-tuning large parameter language models to adapt to specific tasks remains a challenge due to the extensive memory requirements.&#x20;

<mark style="color:blue;">**Quantized Low Rank Adapters (QLORA)**</mark> is an approach that addresses this issue by enabling efficient fine-tuning of quantized LLMs.

{% embed url="<https://github.com/artidoro/qlora>" %}

QLORA <mark style="color:yellow;">**combines quantization techniques with Low Rank Adapters (LoRA)**</mark> to significantly reduce the memory footprint during fine-tuning.&#x20;

The core idea behind QLORA is to perform gradient backpropagation through a 4-bit quantized pretrained language model into the LoRA layers.&#x20;

This approach allows for fine-tuning large models, such as those with 65B parameters, on a single 48GB GPU, a feat that was previously impractical due to the excessive memory requirements of regular 16-bit fine-tuning.

### <mark style="color:purple;">Innovations in QLORA</mark>&#x20;

QLORA introduces several techniques to optimise memory usage without compromising performance:

<mark style="color:green;">**4-bit Normal Float (NF4)**</mark>

NF4 is a new data type that is information-theoretically optimal for representing normally distributed weights. It efficiently captures the statistical properties of the weights, enabling accurate quantization while minimising information loss.

<mark style="color:green;">**Double Quantization**</mark>

QLORA employs a technique called Double Quantization, which quantizes the quantization constants themselves. This additional level of quantization further reduces the memory footprint, resulting in an average saving of about 0.37 bits per parameter.

<mark style="color:green;">**Paged Optimisers**</mark>

To handle memory spikes during the fine-tuning process, QLORA uses Paged Optimisers. These optimisers leverage NVIDIA's unified memory feature to automatically manage memory transfers between the CPU and GPU, preventing out-of-memory errors and ensuring smooth processing even when the GPU memory is exhausted.

### <mark style="color:purple;">Guanaco</mark>

State-of-the-Art Performance Using QLORA, the authors introduce the Guanaco family of models, which achieve state-of-the-art performance on the Vicuna benchmark.&#x20;

The second-best model in the Guanaco family reaches an impressive 97.8% of ChatGPT's performance level.&#x20;

With less than 24 hours of training on a professional GPU, the largest Guanaco model attains 99.3% of ChatGPT's performance, effectively closing the gap. Moreover, the smallest Guanaco model, with only 7B parameters, outperforms the larger 26GB Alpaca model on the same benchmark while requiring just 5GB of memory.

### <mark style="color:purple;">Comprehensive Analysis and Evaluation</mark>&#x20;

QLORA's efficiency enables an extensive analysis of instruction fine-tuning and chatbot performance across various scales, architectures, and datasets.&#x20;

The authors train over 1,000 models, ranging from 80M to 65B parameters, demonstrating QLORA's ability to achieve 16-bit performance and train state-of-the-art chatbots like Guanaco.

The study reveals intriguing insights, such as the importance of data quality over dataset size.&#x20;

A smaller dataset of 9k samples outperforms a much larger 450k sample dataset in chatbot performance, emphasising the significance of dataset suitability for specific tasks.

The evaluation methodology employed in the paper is innovative, using tournament-style benchmarking where models compete against each other, with winners determined by either GPT-4 or human annotators.&#x20;

While GPT-4 and human evaluations largely agree, instances of strong disagreement highlight the uncertainties in model-based evaluation. The qualitative analysis of the Guanaco models provides valuable insights into their successes and failures, complementing the quantitative benchmarks.

### <mark style="color:purple;">Technical Details and Explanations</mark>&#x20;

#### <mark style="color:green;">Block-wise k-bit Quantization</mark>

Quantization is the process of <mark style="color:yellow;">**converting a higher-bit datatype to a lower-bit representation**</mark>.

In QLORA, quantization is performed by rescaling the input data into the target data type's range through normalization. To handle outliers, the input tensor is divided into blocks, and each block is independently quantized with its own quantization constant.

### <mark style="color:purple;">Low-rank Adapters (LoRA) Fine-Tuning</mark>

LoRA is a parameter-efficient fine-tuning method that reduces memory requirements by using a small set of trainable parameters called adapters.&#x20;

These adapters augment the linear projection of the model through an additional factorised projection. LoRA's minimal memory footprint allows for the use of more adapters to enhance performance without significantly increasing the total memory usage.

### <mark style="color:purple;">Gradient Checkpointing</mark>

Gradient checkpointing is a technique used to reduce the memory consumption during the training of deep neural networks.&#x20;

Instead of storing all intermediate activations, gradient checkpointing stores only a subset of them, called checkpoints. During backpropagation, the missing gradients between checkpoints are recomputed on-the-fly, enabling the training of deeper models within the same memory constraints.

### <mark style="color:purple;">4-bit Normal Float (NF4) Quantization</mark>

NF4 quantization is an information-theoretically optimal data type for representing normally distributed weights.&#x20;

It builds upon <mark style="color:blue;">**Quantile Quantization**</mark>, aiming to assign an equal number of values to each quantization bin. &#x20;

The process involves estimating the quantiles of the input tensor, transforming the weights to a fixed distribution, and normalizing them to the range \[-1, 1]. Fast quantile approximation algorithms, such as SRAM quantiles, are employed to handle the computational cost of exact quantile estimation.

### <mark style="color:purple;">Double Quantization (DQ)</mark>

Double Quantization is a technique used to reduce the memory overhead associated with 4-bit quantization.&#x20;

It involves <mark style="color:yellow;">**quantizing the quantization constants**</mark> themselves.&#x20;

First, the quantization constants are quantized (cFP32→cFP8), and then a second quantization step yields quantized constants (cFP8→cFP8) and a second level of quantization constants (cFP32).&#x20;

By using 8-bit Floats with a block size of 256 for the second quantization, the memory footprint per parameter is reduced from 0.5 bits to 0.127 bits, resulting in a significant reduction of 0.373 bits per parameter.

### <mark style="color:purple;">Paged Optimisers with NVIDIA Unified Memory</mark>

Paged Optimisers leverage NVIDIA's unified memory feature to facilitate automatic page-to-page transfers between the CPU and GPU.&#x20;

This feature acts like regular memory paging between CPU RAM and the disk, allowing for error-free GPU processing even when the GPU runs out of memory.&#x20;

Paged memory is allocated for the optimiser states and automatically transferred to CPU RAM when the GPU memory is exhausted, and then paged back into GPU memory when needed.

### <mark style="color:purple;">Comparison of QLORA with Standard Fine-tuning</mark>

The authors compare QLORA with standard full-model fine-tuning to assess whether it can match the performance.&#x20;

Three architectures are considered for the comparison, including various datasets and models of different sizes. The importance of using LoRA on all transformer layers is emphasised, particularly for LLaMA 7B models, to achieve 16-bit performance.&#x20;

### <mark style="color:purple;">Paged Optimizers and Their Performance</mark>

While paged optimisers are crucial for certain tasks, the authors acknowledge the need for future work to characterise potential slowdowns that might occur due to the paging process.&#x20;

### <mark style="color:purple;">Conclusion</mark>&#x20;

QLORA represents a significant advancement in the efficient fine-tuning of large language models. By combining quantization techniques, Low Rank Adapters, and innovative memory management strategies.

QLORA enables the training of state-of-the-art chatbots and achieves performance comparable to ChatGPT while significantly reducing memory requirements.&#x20;

The comprehensive analysis and evaluation conducted in the paper provide valuable insights into the factors influencing model performance, such as data quality and the importance of applying LoRA to all transformer layers.

The introduction of novel techniques like 4-bit Normal Float (NF4) quantization, Double Quantization, and Paged Optimizers demonstrates the potential for further optimization in the fine-tuning process.&#x20;

These innovations not only reduce memory usage but also maintain high performance, making QLORA a promising approach for researchers and practitioners working with large language models.


# QLORA and Fine-Tuning of Quantized Language Models (LMs)

### <mark style="color:purple;">Introduction</mark>&#x20;

Quantized LoRA (QLoRA) is a novel technique introduced by Tim Dettmers and team, addressing the challenges of training large language models.

### <mark style="color:purple;">Low-Rank Adaptation (LoRA) and Matrix Decomposition</mark>&#x20;

At the core of QLoRA is the concept of low-rank adaptation (LoRA), which involves inserting small, low-rank matrices (adapters) between the layers of a pre-trained model.&#x20;

These adapters capture the essential transformations needed for a specific task while keeping the model complexity low.&#x20;

By decomposing the weight matrices into lower-rank representations, LoRA enables efficient fine-tuning of LMs by focusing on updating the adapters instead of the entire model.

### <mark style="color:purple;">Quantization Techniques and Memory Efficiency</mark>&#x20;

QLoRA takes the concept of low-rank adaptation further by introducing quantization techniques to reduce memory usage and computational requirements.&#x20;

The quantization process in QLoRA involves mapping the high-precision weight values to a smaller set of discrete values, reducing the memory footprint of the model.&#x20;

QLoRA is use of a data type called the <mark style="color:blue;">**4-bit normal float (NF4)**</mark>, which is optimised for representing normally distributed weights.&#x20;

Unlike traditional quantization methods that use a single constant for all weights, *<mark style="color:yellow;">**QLoRA computes separate quantization constants for each block of weights**</mark>*, ensuring a more accurate and nuanced quantization.

The quantization constants play a crucial role in this process, serving as scaling factors to map the quantized values back to their original range during dequantization.&#x20;

QLoRA employs a technique called <mark style="color:blue;">**double quantization**</mark>, where both the model weights and the quantization constants themselves are quantized, further reducing memory usage.

### <mark style="color:purple;">Balancing Precision and Efficiency</mark>&#x20;

One of the main challenges in quantizing LLMs is <mark style="color:yellow;">**maintaining the balance between precision and efficiency**</mark>.&#x20;

QLoRA addresses this challenge through the use of the <mark style="color:blue;">**NF4 data type**</mark> and the strategic placement of LoRA adapters throughout the model.&#x20;

The NF4 data type is designed to efficiently use the quantization bins, especially around the centre of the weight distribution where the density of values is highest. This allows for a more accurate representation of the weights while minimizing the loss of precision.

The placement of LoRA adapters is another crucial factor in QLoRA's effectiveness. By distributing the adapters across different layers of the model, QLoRA enables fine-grained control over the model's behaviour at various levels of abstraction.&#x20;

This strategic placement allows for targeted modifications to specific aspects of the model's processing, enhancing the fine-tuning process.

### <mark style="color:purple;">Memory Management and Computational Efficiency</mark>&#x20;

Efficient memory management is a key aspect of QLoRA, particularly when dealing with the activation gradients during training.&#x20;

QLoRA employs techniques such as <mark style="color:blue;">**gradient checkpointing**</mark> and <mark style="color:blue;">**paged optimisers**</mark> to reduce the memory footprint of these gradients and prevent memory spikes.&#x20;

Additionally, QLoRA uses a combination of <mark style="color:blue;">**low-precision storage**</mark> (4-bit) and <mark style="color:blue;">**higher-precision computation**</mark> (16-bit) data types to strike a balance between memory efficiency and computational accuracy.

The choice of data types plays a significant role in QLoRA's performance. The BFloat16 format, commonly used in neural network computations, provides a good balance between precision and efficiency.&#x20;

QLoRA leverages this format for computations, allowing for reduced memory usage while maintaining sufficient precision for effective fine-tuning.

### <mark style="color:purple;">Quantile Quantization and Data Distribution</mark>&#x20;

QLoRA introduces the concept of quantile quantization, which takes into account the statistical properties of the weight distribution.&#x20;

By employing techniques like the NF4 data type, QLoRA ensures that each quantization bin contains an equal number of values from the input tensor, optimising the quantization process for normally distributed weights.&#x20;

This approach leads to a more efficient utilisation of the available quantization bins, particularly in the dense regions of the distribution.

However, quantile quantization also presents challenges, such as the inability to represent zero exactly in the NF4 data type.&#x20;

This can be problematic when dealing with elements like padding, which rely on the presence of true zero values. To address this issue, QLoRA introduces asymmetric data types that allow for an exact zero-point representation, ensuring the accurate handling of zero values in various contexts.

### <mark style="color:purple;">Scaling Behaviour and Model Size</mark>

As LMs continue to grow in size and complexity, understanding their scaling behavior becomes increasingly important.&#x20;

QLoRA sheds light on the peculiar scaling properties observed in large models, where the number of outliers in the weight distribution tends to increase with the model size.  This observation suggests that traditional assumptions about data distributions and quantization strategies may need to be re-evaluated as models scale up.

The relationship between model size and performance is another critical aspect explored in QLoRA.&#x20;

The findings indicate that certain model sizes, such as the 13B parameter range, offer a favourable balance between efficiency and effectiveness. This insight can guide researchers in selecting the optimal model size for specific tasks, considering the trade-offs between computational resources and desired performance.

### <mark style="color:purple;">Hyperparameter Transferability and Model Development</mark>

QLoRA also investigates the transferability of hyperparameters across different model sizes.&#x20;

Surprisingly, the results suggest that hyperparameters optimised for smaller models can be effectively transferred to larger models, reducing the need for extensive tuning at each scale.&#x20;

This finding challenges the conventional wisdom that larger models always require distinct hyperparameter settings, opening up new possibilities for more efficient model development pipelines.

### <mark style="color:purple;">Evaluation Challenges and Future Directions</mark>

Evaluating the performance of LLMs is a complex task, given the lack of standardised benchmarks and the rapidly evolving nature of the field.&#x20;

QLoRA acknowledges these challenges and highlights the need for more comprehensive and widely accepted evaluation protocols. The development of robust and representative benchmarks is crucial for accurately assessing the capabilities of LLMs and comparing different fine-tuning techniques.

Looking ahead, QLoRA presents numerous opportunities for further research and application. The versatility of QLoRA across various domains, such as vision and robotics, suggests its potential as a general-purpose fine-tuning framework.&#x20;

The success of similar approaches, like ControlNet for diffusion models, further validates the effectiveness of low-rank adaptation and quantization techniques in diverse settings.

### <mark style="color:purple;">Conclusion</mark>

Quantized LoRA (QLoRA) represents a significant advancement in the efficient fine-tuning of large language models.&#x20;

By combining low-rank adaptation with quantization techniques, QLoRA enables the effective compression and adaptation of LLMs while maintaining high performance.&#x20;

The strategic placement of LoRA adapters, the use of novel data types like NF4, and the application of quantile quantization contribute to QLoRA's ability to balance precision, efficiency, and scalability.

By addressing the challenges of memory constraints, computational complexity, and model scaling, QLoRA opens up new avenues for research and application, paving the way for more powerful and versatile language models.

However, the journey is far from over. The evaluation challenges, the need for standardized benchmarks, and the exploration of QLoRA's potential across different domains present exciting opportunities for the research community.&#x20;


# ReLoRA: High-Rank Training Through Low-Rank Updates

This <mark style="color:blue;">**December 2023**</mark> paper introduces ReLoRA, a method for efficiently training large neural networks using low-rank updates.&#x20;

The authors argue that despite the current trend of training increasingly large networks with hundreds of billions of parameters, the necessity and theoretical understanding of such overparametrised models remain unclear.&#x20;

ReLoRA aims to address this issue by demonstrating that low-rank updates can be used to train high-rank networks efficiently, potentially challenging the current scaling laws that govern large neural networks.

{% embed url="<https://arxiv.org/abs/2307.05695>" %}
ReLoRA: High-Rank Training Through Low-Rank Updates
{% endembed %}

The paper focuses on applying ReLoRA to pre-training transformer language models with up to 350 million parameters, achieving comparable performance to regular neural network training.&#x20;

The authors suggest that the efficiency of ReLoRA increases with the model size, making it a promising approach for training multi-billion-parameter networks more efficiently.

The paper also discusses the complex relationship between overparametrization and the trainability and generalization of neural networks, referencing concepts such as the Lottery Ticket Hypothesis and parameter-efficient fine-tuning methods like LoRA (Low-Rank Adapters) and Compacter.&#x20;

ReLoRA builds upon these ideas by introducing a method that increases the effective rank of the update in a neural network through restarts, partial optimizer resets, and a jagged-cosine learning rate schedule.

The authors provide a mathematical foundation for ReLoRA, explaining how it expands on the basic idea of LoRA by allowing for multiple restarts, thereby increasing the total rank of the update over time.&#x20;

The paper reports on experiments with transformer language models, emphasizing the efficiency of ReLoRA in terms of both computational resources and training time.

Overall, the paper presents ReLoRA as an innovative approach to efficiently training large-scale neural networks, particularly transformers, by combining low-rank updates with specific training techniques.&#x20;

The authors suggest that their findings could have significant implications for the scaling laws that govern large neural networks and contribute to a better understanding of how to efficiently scale up these models.


# SLoRA: Federated Parameter Efficient Fine-Tuning of Language Models

Leveraging Lora

The paper on S-LoRA focuses on scalable serving of <mark style="color:yellow;">Low-Rank Adaptation (LoRA)</mark> adapters for  language models (LMs).&#x20;

{% embed url="<https://arxiv.org/abs/2308.06522>" %}
SLoRA: Federated Parameter Efficient Fine-Tuning of Language Models
{% endembed %}

### <mark style="color:purple;">**Background and Motivation**</mark>

* <mark style="color:blue;">**Pretrain-then-Finetune Paradigm**</mark><mark style="color:blue;">:</mark> LMs are commonly fine-tuned for specific tasks, leading to numerous fine-tuned variants of a single base model.
* <mark style="color:blue;">**LoRA for Fine-Tuning**</mark><mark style="color:blue;">:</mark> LoRA is a parameter-efficient fine-tuning method that updates only low-rank additive matrices (adapter weights), achieving performance comparable to full-weight fine-tuning.
* <mark style="color:blue;">**Challenges in Serving LoRA Adapters**</mark><mark style="color:blue;">:</mark> While fine-tuning has been extensively researched, the efficient serving of these fine-tuned variants, especially at scale, remains underexplored.

### <mark style="color:purple;">**S-LoRA System**</mark>

* <mark style="color:blue;">**Goal**</mark><mark style="color:blue;">:</mark> To scalably serve thousands of LoRA adapters on a single machine or across multiple GPUs.
* <mark style="color:blue;">**Unified Paging**</mark><mark style="color:blue;">:</mark> S-LoRA introduces Unified Paging to manage dynamic adapter weights with varying ranks and KV cache tensors with different sequence lengths in a unified memory pool.
* <mark style="color:blue;">**Tensor Parallelism and CUDA Kernels**</mark><mark style="color:blue;">:</mark> It employs a tensor parallelism strategy and custom CUDA kernels for heterogeneous batching of LoRA computation.
* <mark style="color:blue;">**Efficiency**</mark>: S-LoRA aims to serve multiple adapters concurrently with high throughput and low latency, significantly improving over traditional methods.

### <mark style="color:purple;">**Technical Innovations**</mark>

* <mark style="color:blue;">**Adapter and Base Model Separation**</mark><mark style="color:blue;">:</mark> S-LoRA separates the batchable base model computation from individual LoRA computations to achieve high-throughput multi-adapter serving.
* <mark style="color:blue;">**Memory Management**</mark><mark style="color:blue;">:</mark> The system stores all adapters in the main memory and fetches the required adapters to the GPU memory as needed, efficiently using GPU memory and reducing fragmentation.
* <mark style="color:blue;">**Serving Throughput and Latency**</mark><mark style="color:blue;">:</mark> Compared to existing libraries (like HuggingFace PEFT and vLLM), S-LoRA can increase throughput by up to 4 times and serve significantly more adapters.

### <mark style="color:purple;">**Potential Impact**</mark>

* <mark style="color:blue;">**Scalable Serving of Fine-Tuned Models**</mark><mark style="color:blue;">:</mark> S-LoRA enables scalable serving of many task-specific fine-tuned models.
* <mark style="color:blue;">**Large-Scale Customized Fine-Tuning Services**</mark><mark style="color:blue;">:</mark> It offers the potential for large-scale customized fine-tuning services, which could be a substantial advancement for deploying personalized AI solutions.

### <mark style="color:purple;">**Challenges in Batched Inference**</mark>

* **Memory Management**: Efficiently managing GPU memory is critical due to the simultaneous serving of multiple LoRA adapters. Storing adapter weights outside the GPU and fetching them dynamically presents challenges in memory fragmentation and I/O overhead.
* **Batching Complexity**: The separated computation of many adapters, each with distinct ranks and stored in non-contiguous memory, complicates the batching process.
* **Multiple GPU Usage**: Utilizing multiple GPUs requires novel parallelism strategies to efficiently handle the added LoRA weights and computations, minimizing communication and memory overheads.

### <mark style="color:purple;">**S-LoRA's Contributions**</mark>

* **Unified Paging**: S-LoRA introduces a unified memory pool to manage dynamic adapter weights and KV cache tensors. This approach reduces memory fragmentation and increases batch size.
* **Heterogeneous Batching**: Custom CUDA kernels are employed for efficiently batching different adapters of varying ranks. These kernels operate on non-contiguous memory, aligning with the memory pool design for efficient batched inference.
* **S-LoRA TP (Tensor Parallelism)**: A novel tensor parallelism strategy ensures effective parallelization across multiple GPUs. It achieves minimal communication cost by scheduling communications on small intermediate tensors and fusing large tensors with the communications of the base model.

### <mark style="color:purple;">**Performance Evaluation**</mark>

* **Model Testing**: S-LoRA was evaluated on models ranging from Llama-7B to Llama-70B.
* **Throughput Enhancement**: Compared to HuggingFace PEFT, a leading parameter-efficient fine-tuning library, S-LoRA can enhance throughput by up to 30 times.
* **Comparison with vLLM**: Against the vLLM system, which naively supports LoRA serving, S-LoRA improves throughput by up to 4 times and significantly increases the number of served adapters.

### <mark style="color:purple;">**Impact and Potential**</mark>

* **Scalability**: S-LoRA's ability to serve thousands of adapters on a single GPU or across multiple GPUs with minimal overhead represents a significant advancement in scalable serving of fine-tuned models.
* **Efficiency in Serving Multiple Tasks**: By overcoming the challenges of memory management, batching complexity, and effective use of multiple GPUs, S-LoRA enables the efficient and scalable serving of a wide range of task-specific fine-tuned models.

In conclusion, the S-LoRA paper presents a comprehensive solution for the scalable serving of LoRA adapters in large language models. By addressing key challenges in memory management, batching, and parallelism, S-LoRA notably enhances throughput and scalability, paving the way for the efficient deployment of highly personalized and diverse AI services.

### <mark style="color:purple;">**LM Architecture and Scale**</mark>

* **Architecture**: Most LMs are based on the transformer architecture.
* **Size**: LMs can have several billion to trillion parameters, resulting in disk sizes from gigabytes to terabytes.
* **Computational and Memory Demands**: The sheer scale of LMs leads to significant computational and memory requirements during serving.

### <mark style="color:purple;">**Inference Process in LMs**</mark>

* **Autoregressive Decoding**: Inference requires iterative autoregressive decoding, where the model <mark style="color:yellow;">encodes the prompt in a forward pass and then decodes the output one token at a time</mark>.
* **Sequential Nature**: The <mark style="color:yellow;">sequential nature of decoding each token, attending to the hidden states of all preceding tokens, makes the process slow</mark>.
* **KV Cache**: There's a need to <mark style="color:yellow;">store the hidden states of all previous tokens</mark>, referred to as the KV cache, which adds substantial memory overhead.

### <mark style="color:purple;">**Challenges in Online Serving**</mark>

* <mark style="color:blue;">**Dynamic Requests**</mark><mark style="color:blue;">:</mark> In online settings, handling requests of varying sequence lengths that arrive dynamically is a challenge.
* <mark style="color:blue;">**Fine-Grained Scheduling**</mark><mark style="color:blue;">:</mark> Orca introduces iteration-level scheduling, batching at the token level instead of the request level, allowing for the addition of new requests to the current batch, thus enhancing throughput.
* <mark style="color:blue;">**vLLM's Memory Optimization**</mark><mark style="color:blue;">:</mark> vLLM optimizes memory efficiency through PagedAttention, which uses concepts from virtual memory and paging to manage KV cache tensors, reducing fragmentation and enabling larger batch sizes.

### <mark style="color:purple;">**Parallelisation Strategies for Large Models**</mark>

* **Necessity for Parallelisation**: Serving models exceeding a single GPU's memory capacity, or meeting strict latency requirements, necessitates parallelization across multiple GPUs.
* **Parallelism Methods**:
  * <mark style="color:blue;">**Tensor Parallelism**</mark><mark style="color:blue;">:</mark> Distributes the computational load of tensor operations across GPUs.
  * <mark style="color:blue;">**Sequence Parallelism**</mark><mark style="color:blue;">:</mark> Focuses on distributing different sequence parts across GPUs.
  * <mark style="color:blue;">**Pipeline Parallelism**</mark><mark style="color:blue;">:</mark> Involves dividing the model into stages, each processed on different GPUs, often in combination with other parallelism methods.

<mark style="color:green;">**Integration of Methods**</mark>

Combining various parallelism techniques (tensor, sequence, pipeline) has been proposed to effectively handle the demands of very large models, ensuring efficient utilization of computational resources and minimizing latency.

In summary, the paper's focus on batching and scheduling highlights S-LoRA's innovative approach to efficiently manage memory and computational resources.&#x20;

This is achieved by separating batched computations for the base model and LoRA adapters, using custom CUDA kernels, and implementing a dynamic iteration-level scheduling strategy.&#x20;

These strategies enable S-LoRA to serve a large number of LoRA adapters concurrently, addressing the challenges of high-throughput and online serving in large language models.

### <mark style="color:purple;">**Batching Strategy**</mark>

* **Goal**: To *<mark style="color:yellow;">support online and high-throughput serving of many LoRA adapters simultaneously</mark>*.
* **Separate Batched Computation**: The base model computation and the LoRA adapters are batched separately. The base computation is batched using General Matrix Multiply (GEMM), while LoRA adapters are handled by custom CUDA kernels.
* **Efficiency**: This separation allows for efficient computation without the need for padding, as would be required by traditional batch GEMM kernels due to heterogeneity in sequence lengths and adapter ranks.

### <mark style="color:purple;">**LoRA Adapter Merging**</mark>

* **Original LoRA Method**: In the original LoRA approach, <mark style="color:yellow;">adapter weights were merged into the base model,</mark> <mark style="color:green;">**creating a new model without additional overhead during inference**</mark>.
* **Limitation in Multiple Adapter Context**: Merging weights for multiple adapters leads to multiple weight copies and misses batching opportunities. Additionally, the original method of adding and subtracting LoRA weights on the fly doesn’t support concurrent inference on separate adapters.

### <mark style="color:purple;">**S-LoRA's Approach**</mark>

* <mark style="color:blue;">**On-the-fly LoRA Computation**</mark><mark style="color:blue;">:</mark> S-LoRA computes the LoRA computation (`xAB`) on-the-fly, avoiding weight duplication and enabling batching of the more costly `xW` operation. This results in considerable savings despite increased computation overhead.
* <mark style="color:blue;">**Custom CUDA Kernels**</mark><mark style="color:blue;">:</mark> Custom CUDA kernels are used for the LoRA computation, optimizing for efficiency and avoiding the limitations of padding required by batch GEMM kernels.

### <mark style="color:purple;">**Memory Management and Adapter Storage**</mark>

* <mark style="color:blue;">**Main Memory Storage**</mark><mark style="color:blue;">:</mark> All LoRA adapters are stored in the main memory. Only the adapters needed for the current batch are fetched to the GPU RAM during inference, which manages the GPU memory more effectively.
* <mark style="color:blue;">**Maximizing Adapter Service**</mark><mark style="color:blue;">:</mark> The maximum number of adapters that can be served is bounded by the main memory size. This approach allows for a higher throughput of adapter serving.

### <mark style="color:purple;">**Iteration-Level Scheduling Batching Strategy**</mark>

* **Adopted from Orca**: The iteration-level scheduling batching strategy from Orca is employed, <mark style="color:yellow;">where requests are scheduled at the token level.</mark>
* **Dynamic Request Incorporation**: New requests are immediately incorporated into the running batch if space is available. Requests exit the batch once they reach the maximum number of generated tokens or other stopping criteria.
* **Memory Management Challenges**: This strategy reduces GPU memory usage but introduces new challenges in memory management.

### <mark style="color:purple;">**Memory Management Techniques**</mark>

* **Efficient Handling of Memory**: Techniques for managing memory efficiently in this dynamic and high-throughput context are discussed in Section 5 of the paper.

In summary, the paper's focus on batching and scheduling highlights S-LoRA's innovative approach to efficiently manage memory and computational resources. This is achieved by <mark style="color:yellow;">separating batched computations for the base model and LoRA adapters, utilizing custom CUDA kernels, and implementing a dynamic iteration-level scheduling strategy.</mark>&#x20;

These strategies enable S-LoRA to serve a large number of LoRA adapters concurrently, addressing the challenges of high-throughput and online serving in large language models.

The section on "Prefetching and Overlapping" and "Custom Kernels" in the S-LoRA paper focuses on optimizing the system's efficiency in handling multiple LoRA adapters. Here's a brief summary:

### <mark style="color:purple;">**Prefetching and Overlapping**</mark>

* <mark style="color:blue;">**Challenge**</mark><mark style="color:blue;">:</mark> I/O overhead from loading and offloading adapter weights, especially with numerous or large adapters, introduces latency.
* <mark style="color:blue;">**Dynamic Prediction Mechanism**</mark><mark style="color:blue;">:</mark> While decoding the current batch, the system predicts which adapters will be needed for the next batch based on the current queue.
* <mark style="color:blue;">**Prefetching Strategy**</mark><mark style="color:blue;">: A</mark>dapters predicted to be needed are prefetched and stored in memory, reducing I/O time and latency for adapter swapping.

### <mark style="color:purple;">**Custom Kernels for Heterogeneous LoRA Batching**</mark>

* <mark style="color:blue;">**Non-Contiguous Memory Handling**</mark><mark style="color:blue;">:</mark> Due to the unified memory pool's design, adapter weights are stored in non-contiguous memory.
* <mark style="color:blue;">**Custom CUDA Kernels**</mark><mark style="color:blue;">:</mark> These kernels are implemented to efficiently handle LoRA computations with varying ranks and sequence lengths in a non-contiguous memory layout.
* <mark style="color:blue;">**Multi-size Batched Gather Matrix-Matrix Multiplication (MBGMM)**</mark><mark style="color:blue;">:</mark> This kernel, implemented in Triton with tiling, is used in the prefill stage to handle a sequence of tokens and gather adapter weights of different ranks.
* <mark style="color:blue;">**Multi-size Batched Gather Matrix-Vector Multiplication (MBGMV)**</mark><mark style="color:blue;">:</mark> Used in the decode stage, this kernel is adapted from Punica to support multiple ranks in a batch and more fine-grained memory gathering for a single token.

### <mark style="color:purple;">**Tensor Parallelism in Batched LoRA Inference**</mark>

* **Objective**: To support multi-GPU inference of large transformer models using batched LoRA inference.
* **Why Tensor Parallelism?**: It's widely used because of its simplicity and compatibility with existing systems. It helps reduce memory usage and latency per GPU when serving large models.

### <mark style="color:purple;">**Partition Strategy**</mark>

* <mark style="color:blue;">**Alignment with Base Model**</mark><mark style="color:blue;">:</mark> The strategy aligns the partitioning of inputs and outputs of the added LoRA computation with those of the base model. This minimizes communication costs by reducing unnecessary communications and fusing some communications.
* <mark style="color:blue;">**Illustration with Feed-Forward Module**</mark>: The partition strategy is demonstrated using a 2-layer MLP (feed-forward module). The base model's first weight matrix (W1) is column-partitioned, and the second (W2) is row-partitioned.
* <mark style="color:blue;">**Added LoRA Computation**</mark><mark style="color:blue;">:</mark> The LoRA computation involves additional matrices (A1, B1, A2, B2). A1 and B1 (for the first weight matrix W1) are column-partitioned, while A2 and B2 (for the second weight W2) are row-partitioned and column-partitioned, respectively.

### <mark style="color:purple;">**Communication in Partitioning**</mark>

* <mark style="color:blue;">**Base Model**</mark><mark style="color:blue;">:</mark> An all-reduce operation accumulates the partial sum from distributed devices.
* <mark style="color:blue;">**LoRA Computation**</mark><mark style="color:blue;">:</mark> Involves all-gather operations to collect intermediate results and all-reduce operations to sum up these results.
* <mark style="color:blue;">**Fusion of Operations**</mark><mark style="color:blue;">:</mark> The strategy includes fusing an all-gather operation with the final all-reduce operation, a novel approach in parallelization.

### <mark style="color:purple;">**Adaptation to Self-Attention Layer**</mark>

* <mark style="color:blue;">**Head Dimension Partitioning**</mark><mark style="color:blue;">:</mark> Similar to the Megatron-LM strategy, the head dimension of the self-attention layer is partitioned.
* <mark style="color:blue;">**Projection Weight Matrices**</mark><mark style="color:blue;">:</mark> The query-key-value projection weight matrix is treated like W1, and the output projection weight matrix is like W2 in the feed-forward module example.

### <mark style="color:purple;">**Communication and Memory Cost Analysis**</mark>

* <mark style="color:blue;">**Costs for Base Model and Added LoRA Computation**</mark><mark style="color:blue;">:</mark> The communication cost for the base model involves one all-reduce operation. For the added LoRA computation, it involves three all-gathers and one all-reduce.
* <mark style="color:blue;">**Negligible Additional Cost**</mark><mark style="color:blue;">:</mark> The extra communication cost introduced by LoRA is small compared to the base model's cost, mainly because the rank of the adapter (r) is much smaller than the hidden size (h).
* <mark style="color:blue;">**Optimal Memory Usage**</mark><mark style="color:blue;">:</mark> The strategy is memory-efficient as it partitions all weight matrices across devices without replication.


# GaLora: Memory-Efficient LLM Training by Gradient Low-Rank Projection

This <mark style="color:blue;">**March 2024 paper**</mark> addresses the memory challenges in training large language models (LLMs) and proposes a novel approach called <mark style="color:yellow;">Gradient Low-Rank Projection (GaLore)</mark> to reduce memory usage while maintaining performance.

{% embed url="<https://arxiv.org/abs/2403.03507>" %}
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
{% endembed %}

### <mark style="color:purple;">Memory challenges in LLM training</mark>

* The growing size of weights and optimizer states in LLMs leads to significant memory requirements.
* Pre-training a LLaMA 7B model from scratch with a single batch size requires at least 58 GB of memory, making it infeasible on consumer-level GPUs like NVIDIA RTX 4090 with 24 GB memory.

#### <mark style="color:green;">Limitations of existing memory-reduction approaches</mark>

* Low-rank adaptation (LoRA) reduces trainable parameters and optimizer states by adding a trainable low-rank matrix to the frozen pre-trained weight in each layer.
* LoRA and its variant ReLoRA have limitations, such as underperforming full-rank training, requiring full-rank warm-up, and altering training dynamics.

#### <mark style="color:green;">Gradient Low-Rank Projection (GaLore):</mark>

* GaLore is a training strategy that <mark style="color:yellow;">allows full-parameter learning while being more memory-efficient than common low-rank adaptation methods.</mark>
* The key idea is to <mark style="color:yellow;">leverage the slow-changing low-rank structure of the gradient matrix</mark> G, rather than approximating the weight matrix itself as low-rank.
* GaLore computes two projection matrices P and Q to project the gradient matrix G into a low-rank form P⊤GQ, reducing the memory cost of optimizer states.
* Occasional updates of P and Q (e.g., every 200 iterations) incur minimal amortized additional computational cost.

#### <mark style="color:green;">Memory efficiency of GaLore</mark>

* GaLore is more memory-efficient than LoRA, yielding up to 30% memory reduction during pre-training.
* 8-bit GaLore, combined with 8-bit optimizers and layer-wise weight update techniques, achieves comparable performance to its full-rank counterpart with less than 10% memory cost of optimizer states.

#### <mark style="color:green;">Feasibility of pre-training on consumer GPUs</mark>

* GaLore enables, for the first time, the feasibility of pre-training a LLaMA 7B model from scratch on a single GPU with 24 GB memory (e.g., NVIDIA RTX 4090) without costly memory offloading techniques.
* GaLore keeps low memory throughout the entire training, without requiring full-rank training warm-up like ReLoRA.

#### <mark style="color:green;">Performance in fine-tuning</mark>

* GaLore is used to fine-tune pre-trained LLMs on GLUE benchmarks with comparable or better results than existing low-rank methods.
* When fine-tuning RoBERTa-Base on GLUE tasks with a rank of 4, GaLore outperforms LoRA.

#### <mark style="color:green;">Compatibility and ease of use</mark>

* As a gradient projection method, GaLore is independent of the choice of optimizers and can be easily plugged into existing ones with only two lines of code.
* GaLore works for popular optimizers such as <mark style="color:blue;">AdamW, 8-bit Adam, and Adafactor</mark>, and its performance is insensitive to the very few hyperparameters it introduces.

#### <mark style="color:green;">Theoretical justification and convergence analysis</mark>

* The paper provides theoretical justification for the low-rankness of gradient updates and convergence analysis of GaLore.

### <mark style="color:purple;">Other Methodologies</mark>

The paper discusses several methodologies for memory-efficient optimization in training large language models (LLMs). Here's a simplified comparison and contrast of the key approaches:

#### <mark style="color:green;">Low-Rank Adaptation (LoRA)</mark>

* LoRA reduces memory footprint by maintaining a low-rank weight adaptor for each layer.
* It introduces additional low-rank adaptors (A and B) to the fixed weight matrix (W0).
* LoRA and its variants have limitations, such as underperforming full-rank training and requiring full-rank warm-up.

#### <mark style="color:green;">Subspace Learning</mark>

* Subspace learning optimizes model weights within a low-rank subspace.
* It leverages the finding that learning primarily occurs within a significantly low-dimensional parameter subspace.
* This notion has been widely used in various domains of machine learning.

#### <mark style="color:green;">Projected Gradient Descent (PGD)</mark>

* PGD is a traditional optimization method that studies gradients in the vector space.
* GaLore is related to PGD but considers the specific gradient form that appears in training multi-layer neural networks.
* GaLore proves properties of the gradients in the matrix space, while traditional PGD treats the objective as a general black-box nonlinear function.

#### <mark style="color:green;">Memory-Efficient Optimisation</mark>

* Various methods have been proposed to reduce the memory cost of gradient statistics for adaptive optimization algorithms.
* Adafactor achieves sub-linear memory cost by factorizing the second-order statistics using a row-column outer product.
* Quantization is widely used to reduce the memory cost of optimizer states.
* Fused gradient computation reduces the memory cost of storing weight gradients during training.

### <mark style="color:purple;">Gradient Low-Rank Projection (GaLore)</mark>

* GaLore is a training strategy that allows full-parameter learning while being more memory-efficient than low-rank adaptation methods.
* It leverages the slow-changing low-rank structure of the gradient matrix G.
* GaLore computes projection matrices P and Q to project the gradient matrix G into a low-rank form, reducing the memory cost of optimizer states.
* Unlike LoRA, GaLore explicitly utilizes low-rank updates instead of introducing additional low-rank adaptors, preserving the original training dynamics.
* GaLore operates independently of the optimizers, as they directly receive the low-rank gradients without knowing their full-rank counterparts.

In summary, LoRA and subspace learning focus on optimizing model weights within a low-rank subspace, while GaLore leverages the low-rank structure of the gradient matrix to reduce memory cost. PGD is a traditional optimization method, and memory-efficient optimization techniques like Adafactor and quantization aim to reduce the memory cost of optimizer states. GaLore differs from LoRA by explicitly utilizing low-rank updates and preserving the original training dynamics, making it more memory-efficient while maintaining performance.

### <mark style="color:purple;">GaLore's memory-efficient training techniques</mark>

#### <mark style="color:green;">Composition of Low-Rank Subspaces</mark>

* GaLore allows switching across low-rank subspaces during training to learn full-rank weights without increasing memory footprint.
* The weight updates are accumulated within each subspace, and the projectors (P and Q) are re-initialized when switching to a new subspace.
* The switching frequency (T) becomes a hyperparameter, with a sweet spot existing between too frequent and too infrequent changes.
* The computational overhead induced by SVD for subspace switching is negligible compared to other memory-efficient training techniques.

#### <mark style="color:green;">Memory-Efficient Optimization</mark>

* GaLore significantly reduces the memory cost of optimizers that rely on component-wise gradient statistics, such as Adam.
* By projecting the gradient G into its low-rank form R, Adam's gradient regularizer only needs to track low-rank gradient statistics.
* GaLore can be applied to other optimizers (e.g., Adafactor) with similar update rules and memory requirements for gradient statistics.
* To achieve the best memory-performance trade-off, GaLore uses only one projection matrix (P or Q) based on the dimensions of the weight matrix.
* GaLore requires less memory than LoRA during training, as it does not need to store a separate low-rank factorization.

#### <mark style="color:green;">Combining with Existing Techniques</mark>

* GaLore is compatible with existing memory-efficient optimization techniques, such as 8-bit optimizers and per-layer weight updates.
* 8-bit Adam optimizer maintains 32-bit optimizer performance at a fraction of the original memory footprint, and GaLore can be directly applied to its implementation.
* Per-layer weight updates are adopted in GaLore to further reduce memory footprint by performing weight updates during backpropagation.

#### <mark style="color:green;">Hyperparameters of GaLore</mark>

* GaLore introduces very few additional hyperparameters: rank (r), subspace change frequency (T), and scale factor (α).
* The rank (r) is also present in LoRA, while the subspace change frequency (T) is specific to GaLore.
* The scale factor (α) controls the strength of the low-rank update and does not depend on the rank (r), unlike LoRA's scale factor (α/r).

GaLore's memory-efficient training techniques, such as low-rank subspace composition and memory-efficient optimisation, enable it to learn full-rank weights while significantly reducing memory footprint.&#x20;

GaLore is compatible with e`x`isting optimisation methods, such as 8-bit Adam and per-layer weight updates, further enhancing its memory efficiency. The introduction of very few additional hyperparameters makes GaLore easy to use and tune for optimal performance.

### <mark style="color:purple;">Conclusion</mark>

In conclusion, this paper introduces GaLore, a novel memory-efficient approach for pre-training and fine-tuning large language models.&#x20;

GaLore employs gradient low-rank projection to significantly reduce the memory footprint required for storing model parameters and optimizer states, achieving up to 65.5% memory savings compared to traditional methods.

Extensive experiments on pre-training LLaMA models up to 7 billion parameters and fine-tuning on the GLUE benchmark demonstrate that <mark style="color:yellow;">GaLore maintains comparable performance to full-rank training while utilizing substantially less memory.</mark>&#x20;

Notably, GaLore enables pre-training 7B models within the memory constraints of consumer GPUs like the RTX 4090, facilitating more accessible large model training.

The success of GaLore highlights the potential of gradient low-rank projection techniques for memory-efficient training of large models.&#x20;

Looking ahead, future research can explore applying GaLore to other model architectures, further improving memory efficiency through quantization or specialized projection matrices, and enabling elastic distributed training on consumer hardware.

Ultimately, GaLore represents a promising step towards democratising the training of large language models by reducing the substantial computational resources traditionally required. By making large model training more accessible, GaLore could foster broader innovation and applications in natural language processing and beyond.


# Hyperparameters

Art and science

Hyperparameters are critical settings or configurations that govern the training process of the model but are not directly learned from the data.&#x20;

Unlike model parameters, which are learned automatically during training (e.g., weights and biases), hyperparameters must be set prior to training and can significantly influence the model's performance, efficiency, and ability to generalize to new tasks.

In the fine-tuning phase, hyperparameters play a pivotal role in adapting a pre-trained model to a specific task without extensive retraining from scratch.&#x20;

This includes settings such as the <mark style="color:blue;">learning rate</mark>, which determines the size of the steps the model takes during optimisation; the <mark style="color:blue;">batch size</mark>, which affects the amount of data processed simultaneously and influences training stability and speed; and the <mark style="color:blue;">number of epochs</mark>, defining how many times the entire dataset is passed through the model.

#### <mark style="color:green;">Selecting the right set of hyperparameters is crucial</mark>

A learning rate too high might cause the model to overshoot the optimal solution, while one too low may result in a painfully slow convergence. Similarly, an excessively large batch size could lead to poor generalization, and too few epochs might underfit the model to the training data.

The process of hyperparameter tuning involves experimenting with different combinations of hyperparameters to find the set that yields the best performance on a validation dataset.&#x20;

In summary, hyperparameters are the knobs and dials of fine-tuning LLMs, offering a way to customise the training process to achieve optimal performance for specific tasks.&#x20;

Proper tuning of these hyperparameters is essential for unleashing the full potential of LLMs, enabling them to adapt and excel in a wide range of applications.


# Batch Size

Choosing the right batch size is critical

The batch size determines the number of samples used in each update during training.&#x20;

Smaller batch sizes can lead to noisier gradient updates and require more iterations, while larger batch sizes can provide more stable updates but may require more memory.  You'll need to balance the trade-offs between memory usage and training stability when choosing the batch size for fine-tuning.

The batch size affects the speed and stability of the training process and can help to prevent overfitting by introducing noise and randomness into the gradient estimates.&#x20;

However, larger batch sizes require more memory to store intermediate activations and gradients during training.

#### <mark style="color:green;">Impact on Model Performance and Training Methods</mark>

* Batch size is a crucial hyperparameter that defines the number of samples to process before updating the internal model parameters.
* The choice of batch size can significantly influence the performance of deep learning-based neural networks.
* Different strategies like **batch gradient descent** (using all training samples), **mini batch** (using a subset of the training data), or **stochastic gradient descent** (updating after every sample) are employed, and *each has a different impact on the learning process*.

#### <mark style="color:green;">Influence on Generalization and Network Behaviour</mark>

* While accuracy is a vital performance metric, generalization—how well a model performs on unseen data—is equally important.
* Larger batch sizes have been observed to lead to poorer network generalization. This is explored in the paper below:

Choosing the appropriate batch size to minimize resource consumption involves a delicate balance between upfront investment and ongoing usage costs.&#x20;

### <mark style="color:purple;">Resource costs associated with increasing batch size can be bifurcated into:</mark>&#x20;

<mark style="color:green;">**Upfront Costs**</mark>

These include expenses incurred for hardware upgrades or developing infrastructure for multi-GPU training. These costs are one-time investments aimed at enhancing computational capacity.

<mark style="color:green;">**Usage Costs**</mark>

These are recurring expenses linked to resource consumption, including cloud provider fees, electricity, and maintenance costs.

Before ramping up the batch size, especially in the initial stages of a project, it's crucial to evaluate the *cost-benefit trade-off of such investments*.&#x20;

While upfront costs might be significant, they can be justified if they lead to considerable reductions in training time and expedite the experimental tuning phase. However, initiating with a simpler training pipeline is advisable to avoid the complexities and potential bugs associated with parallel training setups.

## <mark style="color:purple;">Resource consumption can be quantified as:</mark>

#### <mark style="color:yellow;">**Resource consumption = (resource consumption per step) x (total number of steps)**</mark>

&#x20;Increasing batch size generally leads to a reduction in the total number of training steps. However, its impact on resource consumption is dependent on how it affects the consumption per step:

* If larger batch sizes can be accommodated by existing hardware with a marginal increase in time per step, the overall resource consumption per step might be offset by the decrease in total steps.
* Doubling the batch size might not affect resource consumption if it halves the number of steps required but also doubles hardware usage, maintaining the same total consumption in terms of GPU-hours.
* Should larger batch sizes necessitate hardware upgrades, the spike in consumption per step might surpass the savings achieved from reduced step counts.

Expanding upon these considerations, here are three novel ideas to enhance resource efficiency:

1. <mark style="color:green;">**Predictive Resource Allocation:**</mark> Develop algorithms that can predict optimal resource allocation based on batch size, model complexity, and historical training data. This would help in preemptively adjusting resource usage to minimize costs.
2. <mark style="color:green;">**Dynamic Resource Scaling:**</mark> Implement a system that dynamically scales resources up or down based on real-time training performance metrics. This would ensure that resources are utilized efficiently, scaling up for computationally intensive tasks and scaling down when demand is low.
3. <mark style="color:green;">**Eco-Friendly Scheduling:**</mark> Design a scheduling system that aligns resource-intensive training tasks with times of lower electricity rates or when renewable energy sources are readily available, reducing the environmental impact and operational costs.

In conclusion, while increasing batch size can lead to reduced training steps, its impact on overall resource consumption is multifaceted.&#x20;


# Padding Tokens

### <mark style="color:purple;">What are padding tokens?</mark>

Padding tokens are special tokens added to sequences in a batch to make them *all have the same length.*&#x20;

In natural language processing (NLP) and machine learning tasks, input data often consists of variable-length sequences, such as sentences or phrases. However, most deep learning models, like recurrent neural networks (RNNs) and transformers, <mark style="color:yellow;">require fixed-length input sequences to process data efficiently.</mark>

Padding tokens help address this issue by filling the shorter sequences with non-informative tokens until they match the length of the longest sequence in the batch. These padding tokens usually have no semantic meaning and are ignored by the model during training and inference.1

By *<mark style="color:yellow;">padding the shorter sequences with non-informative tokens</mark>*, the input data can be represented as a fixed-size tensor that the model can process efficiently. This fixed-size tensor is usually a matrix or a higher-dimensional tensor with consistent dimensions across all samples in the batch.

Moreover, using padding tokens helps maintain the structure and readability of the input data. Since padding tokens have no semantic meaning and are ignored during processing, they don't interfere with the model's understanding of the actual content.

In summary, padding is done to facilitate efficient parallel processing of input samples in deep learning models by ensuring that input tensors have consistent dimensions, rather than to satisfy any specific matrix symmetry requirements.

### <mark style="color:purple;">Paddings across tokenization techniques</mark>

&#x20;Therefore, when dealing with sequences, whether they are tokenized at the character level, word level, subword level, or using Byte Pair Encoding (BPE), it is often necessary to ensure that the sequences fed into a model are of consistent length.

Here's how paddings relate to the different levels of tokenization:

#### <mark style="color:green;">Character Level</mark>

Sequences of characters can vary considerably in length. For instance, the word "apple" has 5 characters, while "characterization" has 15.

In tasks like character-level sequence-to-sequence modeling (e.g., for machine translation or spelling correction), sequences might be tokenized into characters, and then padding is used to ensure input and output sequences have consistent lengths.

#### <mark style="color:green;">Word Level</mark>

In models like RNNs, LSTM, or GRU that operate at the word level, sentences or documents are tokenized into words. Since sentences can vary greatly in length, padding is used to ensure all sequences in a batch have the same length.

For instance, "I love apples." might be tokenized as \["I", "love", "apples", "."], but to fit this in a batch with longer sentences, padding tokens might be added.

#### <mark style="color:green;">Subword Level</mark>

Sometimes words are broken down into smaller units, which are not as small as characters but not as long as full words. This is useful for languages or texts with many compound words or morphological variations.

After tokenization, sequences of subwords might still vary in length, requiring padding for consistent model input.

#### <mark style="color:green;">BPE (Byte Pair Encoding)</mark>

BPE is a type of subword tokenization. It starts with character-level tokenization and then iteratively merges the most frequent pairs of characters/subwords until a certain vocabulary size is reached.

Even after BPE tokenization, sentences will result in sequences of varying lengths, hence the need for padding.

### <mark style="color:purple;">How is Padding Typically Done?</mark>

A special token (often "\<PAD>" or just 0 when dealing with integer representations of tokens) is introduced.

Sequences shorter than a specified maximum length are filled with this padding token up to that length.

In deep learning frameworks like TensorFlow and PyTorch, this is usually done using built-in functions (pad\_sequences in TensorFlow's Keras API, or pad\_packed\_sequence in PyTorch).

It's important to note that during model training or evaluation, these padding tokens need to be masked or ignored to ensure they don't influence the model's computations and outputs.

Note: While padding ensures consistency in sequence lengths, it's essential to handle it properly. For many deep learning models, especially transformers, attention masks are used in tandem with padding to ensure the model doesn't "pay attention" to padding tokens.

Here's a simple example to illustrate the concept of padding tokens:

Original sequences (variable length):

\["I", "love", "NLP"]

\["This", "is", "a", "great", "example"]

After adding padding tokens (fixed length):

\["I", "love", "NLP", "\<PAD>", "\<PAD>"]

\["This", "is", "a", "great", "example"]

In this example, the padding token is represented by "\<PAD>". It's added to the first sequence to match the length of the second sequence, which is the longest in the batch.

When using padding tokens, it's *crucial to also use attention masks* or other mechanisms to *inform the model which tokens are actual content and which ones are padding*. This ensures that the model doesn't process the padding tokens and that they don't contribute to the model's predictions or loss calculation.

Padding is done primarily to *<mark style="color:yellow;">facilitate efficient batch processing</mark>* in deep learning models.&#x20;

While the reason is not specifically related to the requirement of matrices being symmetrical, it has to do with the need for input tensors to have consistent dimensions when being processed in parallel.


# Mixed precision training

This <mark style="color:blue;">February 2018</mark> paper introduces a methodology for training deep neural networks using half-precision (FP16) floating point numbers without sacrificing model accuracy or requiring hyperparameter modifications.&#x20;

{% embed url="<https://arxiv.org/abs/1710.03740>" %}
Mixed precision training
{% endembed %}

The authors propose three key techniques to overcome challenges associated with the reduced precision format:

<mark style="color:green;">Maintaining an FP32 master copy of weights</mark> that accumulates gradients after each optimizer step. This master copy is rounded to FP16 for forward and backward passes.

<mark style="color:green;">**Loss scaling to preserve small gradient value**</mark><mark style="color:green;">s</mark> that would otherwise be lost due to the limited range of FP16. Scaling the loss value prior to backpropagation shifts relevant gradients into the representable range.

<mark style="color:green;">**Accumulating FP16 products into FP32 for certain arithmetic operations**</mark> like dot products, while performing others in FP16. This maintains fidelity in crucial network calculations.

### <mark style="color:purple;">Key observations and results</mark>

* FP16 training matches FP32 accuracy with no hyperparameter tuning in most cases. Loss scaling is needed for some models like SSD, machine translation, and language modelling to preserve small gradients.
* FP16 reduces memory consumption and arithmetic time compared to FP32, enabling 2-6x speedups on bandwidth-limited operations on Volta GPUs with tensor cores.
* The FP32 master copy of weights is crucial for convergence, without which models like DeepSpeech 2 Mandarin suffer an 80% relative accuracy loss.
* Speech recognition experiments are the largest models trained, with up to 215M parameters. Interestingly, FP16 slightly outperforms FP32 (5-10%) on these tasks, possibly due to a regularization effect.

Overall, this work demonstrates that reduced precision is a viable approach for accelerating DNN training across a variety of domains without compromising model quality.&#x20;

It overcomes many of the pitfalls of previous attempts at FP16 training. &#x20;

The techniques are straightforward to implement and exhibit promising results on modern tensor core hardware.&#x20;

This sets the stage for wider adoption of reduced precision training to make more efficient use of computational resources and potentially open up new frontiers in deep learning research.


# FP8 Formats for Deep Learning

This <mark style="color:blue;">**September 2022**</mark> paper is a collaborative work by researchers from NVIDIA, Arm, and Intel.&#x20;

The authors propose an <mark style="color:blue;">**8-bit floating-point (FP8) binary interchange format**</mark> for deep learning training and inference, aiming to reduce computational requirements while maintaining result quality.

The paper presents a comprehensive study of the proposed <mark style="color:yellow;">FP8 format for deep learning training and inference</mark>.&#x20;

The authors demonstrate that <mark style="color:yellow;">FP8 can effectively match the result quality of 16-bit training sessions</mark> across a wide range of tasks, model architectures, and sizes, without changing hyperparameters.&#x20;

{% embed url="<https://arxiv.org/abs/2209.05433>" %}
8-bit floating-point (FP8)
{% endembed %}

The authors demonstrate the effectiveness of the FP8 format on various image and language tasks, covering modern neural network architectures such as CNNs, RNNs, and Transformer-based models.&#x20;

They show that *<mark style="color:yellow;">**FP8 training can effectively match the result quality achieved by 16-bit training sessions without changing any hyperparameters**</mark>*.&#x20;

The study includes large language models with up to 175 billion parameters.

The paper also examines FP8 post-training-quantization of language models trained using 16-bit formats that resisted fixed-point int8 quantization.

<figure><img src="/files/8fz3qQDJGlOOonAZnwIC" alt=""><figcaption><p>FP8 is a natural progression for accelerating deep learning (DL) training beyond the 16-bit formats common in modern processors. DL applications require two 8-bit floating point (FP8) binary interchange formats, both supported by Hopper and Ada GPU architectures: E4M3 and E5M2. These types enable doubling the math throughput as well as reducing bandwidth pressure in half; however, their use requires some care due to their narrower range and lower precision compared to the 16-bit formats. We'll cover three aspects of FP8 for deep learning:</p></figcaption></figure>

The presentation below is from NVIDIA's March 2023 NTC Session:

{% file src="/files/eXYQ7ZmdGeGeEAr7cYjy" %}
Dusan Stosic, DL Architecture, NVIDIA Paulius Micikevicius, DL Architecture, NVIDIA
{% endfile %}

### <mark style="color:purple;">Key points from the paper</mark>

1. Reduced precision representation of numbers has been important for deep learning training and inference acceleration.
2. Common floating-point types for training include IEEE single precision, TF32, IEEE half precision, and bfloat16.
3. For inference, fixed-point int8 representation is popular, but it can encounter challenges in maintaining the required accuracy for some applications.
4. The authors propose an FP8 binary format with two encodings: E4M3 and E5M2.
5. The effectiveness of the FP8 format is demonstrated on various image and language tasks, covering modern neural network architectures.
6. FP8 training matches FP16 or bfloat16 training results without changing any model or optimizer hyperparameters.
7. The study includes the training of very large language models, up to 175B parameters.
8. FP8 post-training-quantization is examined for language models trained using 16-bit formats that resisted fixed-point int8 quantization.

The paper discusses several technical aspects of using FP8 formats in deep learning.&#x20;

#### <mark style="color:green;">Precision of mathematical operations</mark>

* When performing mathematical operations on FP8 inputs, the outputs are usually produced in a higher precision format, such as single-precision floating-point (FP32).
* This is similar to how operations on 16-bit floating-point formats (FP16 and bfloat16) are handled in current CPUs, GPUs, and TPUs.
* For example, matrix multiplication or dot-product instructions produce FP32 outputs, while simpler operations like nonlinearities or normalizations are performed after casting the FP8 inputs to FP32.

#### <mark style="color:green;">Scaling factors</mark>

* To better utilise the limited range of FP8 formats, higher-precision values need to be multiplied by a scaling factor before being cast to FP8.
* This process is similar to the loss-scaling technique used in mixed-precision training with FP16, where gradients are scaled to fit within the FP16 range.
* Some networks may require per-tensor scaling factors because the FP8 dynamic range is not sufficient to cover the entire range of important values across all tensors.
* The general idea is to choose a scaling factor that brings the maximum magnitude in the tensor close to the maximum representable magnitude in the corresponding FP8 format.
* Values that overflow are then saturated to the maximum representable value.

#### <mark style="color:green;">Unscaling</mark>

* After converting FP8 values back to a higher precision or after performing arithmetic instructions that produce a higher-precision output, the values need to be unscaled by multiplying them with the inverse of the scaling factor.
* This requires only a minimal amount of additional arithmetic and is amortized over many multiply-accumulate operations with FP8 inputs.

In summary, the technical aspects discussed in the paper focus on the precision of mathematical operations, the use of scaling factors, unscaling, type conversion, and the specific details of the FP8 formats.&#x20;

These considerations are crucial for effectively using FP8 in deep learning while maintaining accuracy and performance.




---

[Next Page](/llms-full.txt/1)

