An enhanced text compression approach using transformer-based language models
| bracu.type.group | Research Publications | |
| datacite.rights | Open Access | |
| dc.contributor.author | Rahman C.M. | |
| dc.contributor.author | Sobhani M.E. | |
| dc.contributor.author | Rodela A.T. | |
| dc.contributor.author | Shatabda, Swakkhar | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-09-05T17:50:08Z | |
| dc.date.available | 2026-09-05T17:50:08Z | |
| dc.date.issued | 2024-01-01 | |
| dc.description.abstract | Text compression shrinks textual data while keeping crucial information, eradicating constraints on storage, band-width, and computational efficacy. The integration of lossless compression techniques with transformer-based text decompression has received negligible attention, despite the increasing volume of English text data in communication. The primary barrier in advancing text compression and restoration involves optimizing transformer-based approaches with efficient pre-processing and integrating lossless compression algorithms, that remained unresolved in the prior attempts. Here, we propose a transformer-based method named RejuvenateFormer for text decompression, addressing prior issues by harnessing a new pre-processing technique and a lossless compression method. Our meticulous pre-processing technique incorporating the Lempel-Ziv-Welch algorithm achieves compression ratios of 12.57, 13.38, and 11.42 on the BookCorpus, EN-DE, and EN-FR corpora, thus showing state-of-the-art compression ratios compared to other deep learning and traditional approaches. Furthermore, the RejuvenateFormer achieves a BLEU score of 27.31, 25.78, and 50.45 on the EN-DE, EN-FR, and BookCorpus corpora, showcasing its comprehensive efficacy. In contrast, the pre-trained T5-Small exhibits better performance over prior state-of-the-art models. | |
| dc.description.version | Published | |
| dc.format.extent | 6 pages | |
| dc.identifier.citation | C. M. Rahman, M. E. Sobhani, A. T. Rodela and S. Shatabda, "An Enhanced Text Compression Approach Using Transformer-based Language Models," 2024 IEEE Region 10 Symposium (TENSYMP), New Delhi, India, 2024, pp. 1-6, doi: 10.1109/TENSYMP61132.2024.10752239. | |
| dc.identifier.doi | 10.1109/TENSYMP61132.2024.10752239 | |
| dc.identifier.issn | 9798350364866 | |
| dc.identifier.other | 2-s2.0-85211953970 | |
| dc.identifier.uri | https://hdl.handle.net/10361/29754 | |
| dc.language.iso | en_US | |
| dc.publisher | Institute of Electrical and Electronics Engineers Inc. | |
| dc.relation.hasversion | 10.1109/TENSYMP61132.2024.10752239 | |
| dc.relation.ispartof | 2024 IEEE Region 10 Symposium Tensymp 2024 | |
| dc.relation.ispartofseries | 2024 IEEE Region 10 Symposium Tensymp 2024 | |
| dc.relation.uri | https://ieeexplore.ieee.org/document/10752239 | |
| dc.subject | Compression ratio | |
| dc.subject | Deep learning | |
| dc.subject | Lossless compression | |
| dc.subject | Lossy compression | |
| dc.subject | Text compression | |
| dc.subject | Transformer | |
| dc.subject.lcsh | Electric transformers. | |
| dc.subject.lcsh | Machine learning. | |
| dc.title | An enhanced text compression approach using transformer-based language models | |
| dc.type | Conference Proceeding | |
| person.affiliation.name | State University of Bangladesh | |
| person.affiliation.name | United International University | |
| person.affiliation.name | United International University | |
| person.affiliation.name | BRAC University | |
| person.identifier.scopus-author-id | 60355011900 | |
| person.identifier.scopus-author-id | 58886631100 | |
| person.identifier.scopus-author-id | 58661115200 | |
| person.identifier.scopus-author-id | 56037035700 |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- An_Enhanced_Text_Compression_Approach_Using_Transformer-based_Language_Models.pdf
- Size:
- 369.64 KB
- Format:
- Adobe Portable Document Format
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description: