---
title: LLM Performance and Acceleration on PowerVR - Part 2
description: Discover how Imagination's PowerVR GPUs accelerate large language model inferencing with Llama.cpp, enhancing performance and efficiency for various applications.
image: https://blog.imaginationtech.com/hubfs/AdobeStock_1830810279.jpeg
---

[![Imagination](https://blog.imaginationtech.com/hubfs/Imagination_December2019%20Theme/Images/Header%20Logo%20White.svg) ![Imagination](https://blog.imaginationtech.com/hubfs/Imagination_December2019%20Theme/Images/Header%20Logo%20Black.svg)](https://blog.imaginationtech.com/)

- [Products & Solutions](https://www.imaginationtech.com/products/) 
    - [GPU](https://www.imaginationtech.com/products/gpu/)
    - [CPU](https://www.imaginationtech.com/products/cpu/)
    - [AI](https://www.imaginationtech.com/products/ai/)
    - [Ethernet](https://www.imaginationtech.com/products/ethernet/)
    - [Ray Tracing](https://www.imaginationtech.com/products/ray-tracing/)
    - [Automotive](https://www.imaginationtech.com/products/automotive/)
    - [Consumer](https://www.imaginationtech.com/products/consumer/)
    - [Desktop](https://www.imaginationtech.com/products/desktop/)
    - [Mobile](https://www.imaginationtech.com/products/mobile/)
    - [Product Finder](https://www.imaginationtech.com/products/catalog/)
- [Developers](https://developer.imaginationtech.com/)
- [News & Insights](https://www.imaginationtech.com/news/) 
    - [Newsroom](https://www.imaginationtech.com/news/?query=news&filter-language=english&filter-search=)
    - [Resources Hub](https://www.imaginationtech.com/resources/?query=resource&filter-language=english&filter-search=)
    - [Blog](https://blog.imaginationtech.com/)
    - [Events](https://www.imaginationtech.com/events/)
    - [Webinars](https://www.imaginationtech.com/webinars/)
    - [University Programme](https://university.imgtec.com/)
    - [中文社区](http://imgtec.eetrend.com/)
- [Company](https://www.imaginationtech.com/about/) 
    - [About Us](https://www.imaginationtech.com/about/)
    - [Diversity, Equity and Inclusion](https://www.imaginationtech.com/diversity-equity-and-inclusion/)
    - [Imagining a Sustainable Future](https://www.imaginationtech.com/imagining-a-sustainable-future/)
    - [Corporate Social Responsibility](https://www.imaginationtech.com/csr/)
    - [Leadership Team](https://www.imaginationtech.com/leadership/)
    - [Contact Us](https://www.imaginationtech.com/contact-us/)
- [Careers](https://www.imaginationtech.com/careers/) 
    - [Work at Imagination](https://www.imaginationtech.com/careers/)
    - [Life at Imagination](https://www.imaginationtech.com/careers/life-at-imagination/)
    - [Hybrid Working](https://www.imaginationtech.com/careers/hybrid-working/)
    - [Early Careers](https://www.imaginationtech.com/careers/early-careers/)
    - [Search Jobs](https://www.imaginationtech.com/careers/vacancies/)

- [Products & Solutions](https://www.imaginationtech.com/products/) 
    - [GPU](https://www.imaginationtech.com/products/gpu/)
    - [CPU](https://www.imaginationtech.com/products/cpu/)
    - [AI](https://www.imaginationtech.com/products/ai/)
    - [Ethernet](https://www.imaginationtech.com/products/ethernet/)
    - [Ray Tracing](https://www.imaginationtech.com/products/ray-tracing/)
    - [Automotive](https://www.imaginationtech.com/products/automotive/)
    - [Consumer](https://www.imaginationtech.com/products/consumer/)
    - [Desktop](https://www.imaginationtech.com/products/desktop/)
    - [Mobile](https://www.imaginationtech.com/products/mobile/)
    - [Product Finder](https://www.imaginationtech.com/products/catalog/)
- [Developers](https://developer.imaginationtech.com/)
- [News & Insights](https://www.imaginationtech.com/news/) 
    - [Newsroom](https://www.imaginationtech.com/news/?query=news&filter-language=english&filter-search=)
    - [Resources Hub](https://www.imaginationtech.com/resources/?query=resource&filter-language=english&filter-search=)
    - [Blog](https://blog.imaginationtech.com/)
    - [Events](https://www.imaginationtech.com/events/)
    - [Webinars](https://www.imaginationtech.com/webinars/)
    - [University Programme](https://university.imgtec.com/)
    - [中文社区](http://imgtec.eetrend.com/)
- [Company](https://www.imaginationtech.com/about/) 
    - [About Us](https://www.imaginationtech.com/about/)
    - [Diversity, Equity and Inclusion](https://www.imaginationtech.com/diversity-equity-and-inclusion/)
    - [Imagining a Sustainable Future](https://www.imaginationtech.com/imagining-a-sustainable-future/)
    - [Corporate Social Responsibility](https://www.imaginationtech.com/csr/)
    - [Leadership Team](https://www.imaginationtech.com/leadership/)
    - [Contact Us](https://www.imaginationtech.com/contact-us/)
- [Careers](https://www.imaginationtech.com/careers/) 
    - [Work at Imagination](https://www.imaginationtech.com/careers/)
    - [Life at Imagination](https://www.imaginationtech.com/careers/life-at-imagination/)
    - [Hybrid Working](https://www.imaginationtech.com/careers/hybrid-working/)
    - [Early Careers](https://www.imaginationtech.com/careers/early-careers/)
    - [Search Jobs](https://www.imaginationtech.com/careers/vacancies/)

×

[AI](https://blog.imaginationtech.com/tag/ai) [GPU](https://blog.imaginationtech.com/tag/gpu)

## LLM Performance and Acceleration on PowerVR - Part 2

![Picture of Alex Pim](https://blog.imaginationtech.com/hs-fs/hubfs/Blog_Asset/AlexPim.jpg?width=60&name=AlexPim.jpg)

 By [Alex Pim](https://blog.imaginationtech.com/author/alex-pim)

 Sep 25, 2026  |  3 min read

- 25 September 2026
- [Alex Pim](https://blog.imaginationtech.com/author/alex-pim)

## Part 2: Recap

Welcome to the second part of this two-part blog series about how to accelerate large language model inferencing on PowerVR GPUs. 

In the first part of this blog, we explained what a large language model is and introduced the concept of attention, as well as token generation using KV caching. We then explained the two modes of operation of an autoregressive large language model, and why these models can be so compute intensive to run.

We then introduced two operational metrics; Time to first token and Inter-token latency, which are commonly used in the industry to describe the inference performance of large language models - so if you missed part one, it can be found here.

In part two we will introduce a popular open-source large language model inference platform called Llama.cpp and how, with the contributions Imagination has made, it can be used to deploy and benchmark large language models running on a PowerVR GPU using the platform’s OpenCL Compute back-end.

### What is Llama.cpp?

Llama.cpp is an open-source implementation of a large language model inference engine and contributions have been made to support the offloading of repeating matrix-matrix, and matrix-vector multiplication-based transformer layers to GPUs, via a range of configurable back-end implementations, including OpenCL and Vulkan.

Llama.cpp has helped move large parameter models, previously constrained to data centres, onto other hardware platforms including edge devices such as mobile phones. The comprehensive software flow from PyTorch model to GGUF format (the internal model format of llama.cpp) provides model benchmarking and high-performance deployments of several LLM architectures, many of which can be found in well-known model databases such as Hugging Face.

A key part of the Llama.cpp platform is the capability to process models that use pre-quantised weights formats like 4-bit integer as well as the commonly used F16 and F32 formats.

Supporting compressed data formats such as Q4\_0 significantly reduces the memory footprint of the model parameters (and thus lowers bandwidth of loading weights into the GPU’s local memory) without compromising the accuracy of the model. 

This is one of the major requirements of moving large language model inferencing from data centre deployment to edge devices like mobile phones or automotive applications.

### How can a PowerVR GPU help with LLM Compute?

It is well known that GPUs, originally designed to offload intensive graphics calculations from the CPU, are also very good at executing highly parallelisable workloads such as matrix-matrix and matrix-vector multiplication, and these operations as discussed previously are at the heart of LLM compute, and many other types of neural networks.

The contributions that Imagination has made to the OpenCL compute back-end in Llama.cpp allow off-the-shelf support for PowerVR GPUs in the form a) supporting the existing compute kernels in the default OpenCL back-end, b) the addition of optimised kernels designed specifically for the PowerVR GPU architecture in the form of the PowerVR optimised OpenCL back-end and c) an evolution of optimised kernels that implements cooperative matrix multiplication; a compute operation now supported on a range of GPUs, including Imagination’s latest E-Series GPU IP.

This latest category of cooperative matrix multiplication kernels will be pushed upstream to the PowerVR optimised OpenCL back-end of Llama.cpp in the coming weeks.

### Default OpenCL Back-end - Running on PowerVR

With a few modifications to the existing OpenCL back-end code to recognise Imagination as GPU vendor and to register a sub-workgroup size of 128 in the kernel code, a PowerVR GPU can be used to unlock a wide range of functionality within Llama.cpp for large language model inferencing.

### Default Vulkan Back-end

As well as supporting OpenCL, PowerVR GPUs also support the Vulkan Compute platform, so with similar modifications to the Vulkan back-end to identify the PowerVR GPU and set the default sub-workgroup size, large language model inferencing can be carried out in Llama.cpp with the Vulkan back-end as well.

### PowerVR Optimised Kernel Code

However, to achieve maximum performance on a PowerVR GPU, Imagination’s performance-tuned kernels have also been added to the OpenCL back-end as a build configurable extension. 

These high-performance kernels add support for F32, F16 and Q4\_0 weights format and use very specific memory access patterns that are tailored to the architecture of PowerVR GPUs. 

### Cooperative Matrix Multiply – Now supported by PowerVR!

If you are lucky enough to have access to PowerVR’s latest GPU, then large language inference performance just keeps getting better and better…

With new hardware support for the “SPV\_KHR\_cooperative\_matrix” Khronos extension, a single core PowerVR E-series GPU can achieve over 500 tokens/sec in prefill mode for a prompt size of 2048 when running the QWEN3.5 4 billion parameter model with F16 weights format.

And if you would like to see an example of an agentic workload running on a PowerVR GPU using Llama.cpp – then check out [this blog](https://blog.imaginationtech.com/running-opencode-with-llama.cpp-and-qwen-3.5-on-e-series?hs_preview=vfBuHwQt-222174085976) on running OpenCode with QWEN 3.5 on E-Series.

### Performance Summary

Here we show the optimised kernel performance on three PowerVR GPUs when inferencing using QWEN-3.5 4 billion parameter model and F16 weights format:  
 

![llama.cpp qwen3.5 performance](https://blog.imaginationtech.com/hs-fs/hubfs/llama.cpp%20qwen3.5%20performance.png?width=4106&height=1772&name=llama.cpp%20qwen3.5%20performance.png)

*Figure 1 - Normalised LLM inference Prefill performance on PowerVR GPUs*

### Final Words

Imagination’s new E-series GPU is over 15% more efficient and 400% more performant for matrix-multiplication than the previous D-series architecture, resulting in a 4.7x performance uplift when inferencing using large language models.

And the contributions being made by Imagination to the OpenCL back-end in the Llama.cpp open-source repository not only allow out-of-the box compute offloading to any PowerVR GPU, but also the option of using Imagination’s performance-tuned kernels to unlock the full potential of a PowerVR GPU.

So, this concludes our two-part blog series introducing LLM Acceleration and Performance on PowerVR GPUs. We hope you found this post useful, and encourage you to have a go at running large language models using Llama.cpp on the latest PowerVR GPUs.

Happy inferencing!

[AI](https://blog.imaginationtech.com/tag/ai) [GPU](https://blog.imaginationtech.com/tag/gpu)

#### Share this post

- <https://twitter.com/intent/tweet?url=https://blog.imaginationtech.com/llm-performance-and-acceleration-on-powervr-part-2&text=LLM%20Performance%20and%20Acceleration%20on%20PowerVR%20-%20Part%202>
- [mailto:?subject=https://blog.imaginationtech.com/llm-performance-and-acceleration-on-powervr-part-2](mailto:?subject=https://blog.imaginationtech.com/llm-performance-and-acceleration-on-powervr-part-2)

##### About the Author

![Picture of Alex Pim](https://blog.imaginationtech.com/hs-fs/hubfs/Blog_Asset/AlexPim.jpg?width=60&name=AlexPim.jpg)

[Alex Pim](https://blog.imaginationtech.com/author/alex-pim)

 Alex joined Imagination in 2002, shortly after gaining an honours degree in Artificial Intelligence & Computer Science from the University of Birmingham. Since then, he has contributed to almost all areas of software engineering within the business including Digital Video, Digital Communications and Graphics Processing Units. With a personal interest in AI, and extensive knowledge of the company’s IP he has previously run the AI Engineering Division of PowerVR and now works in the Software Architecture team as an Engineering Fellow, taking Imagination into a new era of efficient and high performing Compute solutions.

[More from Alex Pim](https://blog.imaginationtech.com/author/alex-pim)

### Subscribe to our blog

### Introducing E-Series GPU

 Imagination E-Series GPU IP breaks through the barriers of throughput, memory, and power at the edge.Built for next-gen AI, E-Series delivers high performance and adaptability across mobile, automotive, desktop, and consumer markets, avoiding the rigidity of fixed-function accelerators while retaining the flexibility of GPU programmability. 

[Learn more](https://www.imaginationtech.com/products/e-series/)

### Read Next

[![](https://blog.imaginationtech.com/hubfs/Untitled%20design%20(5).png)](https://blog.imaginationtech.com/from-graphics-to-ai-and-back-again)

## [From graphics to AI, and back again](https://blog.imaginationtech.com/from-graphics-to-ai-and-back-again)

[GPU](https://blog.imaginationtech.com/tag/gpu)

4 min read

Forty-five years ago, my father was studying radio engineering in the Soviet Union. The transistors...

[Read more](https://blog.imaginationtech.com/from-graphics-to-ai-and-back-again)

[![](https://blog.imaginationtech.com/hubfs/Untitled%20design%20(4).png)](https://blog.imaginationtech.com/running-opencode-with-llama.cpp-and-qwen-3.5-on-e-series)

## [Running OpenCode with llama.cpp and Qwen 3.5 on E-Series](https://blog.imaginationtech.com/running-opencode-with-llama.cpp-and-qwen-3.5-on-e-series)

[AI](https://blog.imaginationtech.com/tag/ai) [Software](https://blog.imaginationtech.com/tag/software)

5 min read

Edge AI gets more interesting when a model does more than answer prompts. It plans, calls tools,...

[Read more](https://blog.imaginationtech.com/running-opencode-with-llama.cpp-and-qwen-3.5-on-e-series)

[![](https://blog.imaginationtech.com/hubfs/AdobeStock_223397823.jpeg)](https://blog.imaginationtech.com/imagination-chiplets-scaling-system-design-beyond-the-monolithic-soc)

## [Imagination & Chiplets: Scaling System Design Beyond the Monolithic SoC](https://blog.imaginationtech.com/imagination-chiplets-scaling-system-design-beyond-the-monolithic-soc)

[Automotive](https://blog.imaginationtech.com/tag/automotive) [GPU](https://blog.imaginationtech.com/tag/gpu)

4 min read

For many years, improvements in SoC performance came from process scaling and integrating more...

[Read more](https://blog.imaginationtech.com/imagination-chiplets-scaling-system-design-beyond-the-monolithic-soc)

- Products 
    - [GPU](https://www.imaginationtech.com/products/gpu/)
    - [CPU](https://www.imaginationtech.com/products/cpu/)
    - [AI](https://www.imaginationtech.com/products/ai/)
    - [Ethernet](https://www.imaginationtech.com/products/ethernet/)
    - [Ray Tracing](https://www.imaginationtech.com/products/ray-tracing/)
    - [Open Access](https://www.imaginationtech.com/products/open-access/)
    - [Design Optimization Kit](https://www.imaginationtech.com/products/design-optimisation-kit/)
    - [Product Finder](https://www.imaginationtech.com/products/catalog/)
- Developers 
    - [Developers](https://developer.imaginationtech.com/)
    - [PowerVR SDK and Tools](https://developer.imaginationtech.com/powervr-sdk/)
    - [Developer Downloads](https://developer.imaginationtech.com/downloads/)
    - [Developer Documentation](https://docs.imgtec.com/)
    - [Developer Forums](https://forums.imgtec.com)
- About Imagination 
    - [About Us](https://www.imaginationtech.com/about/)
    - [Career Opportunities](https://www.imaginationtech.com/careers/)
    - [Imagining a Sustainable Future](https://www.imaginationtech.com/imagining-a-sustainable-future/)
    - [Corporate Social Responsibility](https://www.imaginationtech.com/csr/)
    - [Imagination Leadership](https://www.imaginationtech.com/leadership/)
    - [Contact Imagination](https://www.imaginationtech.com/contact-us/)
- Applications 
    - [Automotive](https://www.imaginationtech.com/products/automotive/)
    - [Consumer](https://www.imaginationtech.com/products/consumer/)
    - [Desktop](https://www.imaginationtech.com/products/desktop/)
    - [Mobile](https://www.imaginationtech.com/products/mobile/)
- News & Insights 
    - [News](https://www.imaginationtech.com/news/)
    - [Resources](https://www.imaginationtech.com/resources/?query=resource&filter-language=english&filter-search=)
    - [Blog](https://blog.imaginationtech.com/)
    - [Events](https://www.imaginationtech.com/events/)
    - [Webinars](https://www.imaginationtech.com/webinars/)
    - [The Future of Automotive](https://www.imaginationtech.com/future-of-automotive/)
    - [University Programme](https://university.imgtec.com/)

[linkedin](https://www.linkedin.com/company/imgtec) [twitter](https://twitter.com/imaginationtech) [facebook](https://www.facebook.com/imgtec/) [youtube](https://www.youtube.com/user/Imgtec/) [github](https://github.com/powervr-graphics)

![Imagination](https://blog.imaginationtech.com/hubfs/Imagination_December2019%20Theme/Images/Footer%20Logo.svg)

© Imagination Technologies Limited. All rights reserved.

- [Privacy Policy](https://www.imaginationtech.com/privacy/)
- [Use of Cookies](https://www.imaginationtech.com/cookies/)
- [Terms of Use](https://www.imaginationtech.com/terms/)
- [Trademarks](https://www.imaginationtech.com/trademarks/)
- [Quality Policy](https://www.imaginationtech.com/quality-policy/)

[linkedin](https://www.linkedin.com/company/imgtec) [twitter](https://twitter.com/imaginationtech) [facebook](https://www.facebook.com/imgtec/) [youtube](https://www.youtube.com/user/Imgtec/) [github](https://github.com/powervr-graphics)

```json
{
  "@context" : "https://schema.org",
  "@type" : "VideoObject",
  "contentUrl" : "https://2426966.fs1.hubspotusercontent-na1.net/hubfs/2426966/video_assets/222365834860/inherited/web_optimized.mp4",
  "dateModified" : "2026-09-21T11:10:58.274Z",
  "duration" : "PT1M19S",
  "height" : 1080,
  "name" : "Neural Super Resolution - Public",
  "thumbnailUrl" : "https://resources.imaginationtech.com/hubfs/Social/Neural%20Super%20Resolution%20-%20Public.mp4/medium.jpg",
  "uploadDate" : "2026-09-21T11:10:47.048Z",
  "width" : 1920
}
```