Research topic

LLM and Transformer Efficiency Research

This topic covers methods for reducing the computation, memory, energy, or deployment cost of Transformer and large-language-model workloads. The records identify the efficiency intervention, model or workload, quality metric, resource measure, and trade-off reported by the source.

Related search terms: large-language-model efficiency · Transformer efficiency · efficient AI systems

2 papers in this topic, ordered newest first. The detailed paper records contain the evidence-grounded methods, tools, datasets, findings, and citation guidance.

Selected papers

2025 · 2025 IEEE International Conference on Collaborative Advances in Software and COmputiNg (CASCON)

Energy Consumption Analysis of Large Language Models Across CPU and GPU Using Diverse Metric Types

Tong Zhang, Leila Tahmooresnejad, Naser Ezzati-Jivan

The paper models LLM inference energy separately on CPU and GPU using hardware counters, device metrics, and task/model features, then compares classical and neural regressors across language tasks.

Keywords: LLM energy consumption · CPU energy · GPU energy · green AI · static metrics

Read the detailed paper record · · Authoritative source

2025 · ACM/SPEC International Conference on Performance Engineering (ICPE)

Optimization Strategies for Enhancing Resource Efficiency in Transformers & Large Language Models

Tom Wallace, Beatrice M. Ombuki-Berman, Naser Ezzati-Jivan

The paper compares compression and optimization strategies for reducing the resource cost of Transformer and large-language-model workloads while retaining useful accuracy.

Keywords: transformers · quantization · knowledge distillation · pruning · 4-bit quantization

Read the detailed paper record · · Authoritative source

Related topics