This topic covers methods for reducing the computation, memory, energy, or deployment cost of Transformer and large-language-model workloads. The records identify the efficiency intervention, model or workload, quality metric, resource measure, and trade-off reported by the source.
Related search terms: large-language-model efficiency · Transformer efficiency · efficient AI systems
2 papers in this topic, ordered newest first. The detailed paper records contain the evidence-grounded methods, tools, datasets, findings, and citation guidance.
Selected papers
2025 · 2025 IEEE International Conference on Collaborative Advances in Software and COmputiNg (CASCON)
The paper models LLM inference energy separately on CPU and GPU using hardware counters, device metrics, and task/model features, then compares classical and neural regressors across language tasks.
Keywords: LLM energy consumption · CPU energy · GPU energy · green AI · static metrics
Tom Wallace, Beatrice M. Ombuki-Berman, Naser Ezzati-Jivan
The paper compares compression and optimization strategies for reducing the resource cost of Transformer and large-language-model workloads while retaining useful accuracy.