View 1 Understanding Cryptography 2024 flipbook.
Understanding Cryptography From Established Symmetric and Asymmetric Ciphers to Post-Quantum Algorithms Second Edition Christof Paar · Jan Pelzl · Tim Güneysu
Understanding Cryptography
Christof Paar • Jan Pelzl • Tim Güneysu Understanding Cryptography Second Edition From Established Symmetric and Asymmetric Ciphers to Post-Quantum Algorithms
ISBN 978-3-662-69006-2 ISBN 978-3-662-69007-9 (eBook) https://doi.org/10.1007/978-3-662-69007-9 Originally published under: Paar, C. and Pelzl, J. This work is subject to copyright. All rights are solely and exclusively licensed by the Publisher, whether the whole or part of the material is concerned, specifically the rights of translation, reprinting, reuse of illustrations, recitation, broadcasting, reproduction on microfilms or in any other physical way, and transmission or information storage and retrieval, electronic adaptation, computer software, or by similar or dissimilar methodology now known or hereafter developed. The use of general descriptive names, registered names, trademarks, service marks, etc. in this publication does not imply, even in the absence of a specific statement, that such names are exempt from the relevant protective laws and regulations and therefore free for general use. The publisher, the authors, and the editors are safe to assume that the advice and information in this book are believed to be true and accurate at the date of publication. Neither the publisher nor the authors or the editors give a warranty, expressed or implied, with respect to the material contained herein or for any errors or omissions that may have been made. The publisher remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. This Springer imprint is published by the registered company Springer-Verlag GmbH, DE, part of Springer Nature. The registered company address is: Heidelberger Platz 3, 14197 Berlin, Germany 1st edition: © Springer-Verlag Berlin Heidelberg 2010 2nd edition: © The Editor(s) (if applicable) and The Author(s), under exclusive license to Springer- Verlag GmbH, DE, part of Springer Nature 2024 Christof Paar Max Planck Institute for Security and Privacy Bochum, Germany Jan Pelzl Hamm-Lippstadt University of Applied Sciences Hamm, Germany Tim Güneysu Ruhr University Bochum Bochum, Germany If disposing of this product, please recycle the paper.
To Flo, Maja, Noah and Sarah to Greta, Karl, Thea, Klemens and Nele as well as to Elisa, Benno and Sindy
Foreword Cryptography is a critical component of today’s information infrastructure; it is what enables distributed information systems to exist and to work properly. Without it, users would not be able to securely authenticate themselves to websites, secure communications wouldn’t exist, and privacy would be unachievable. Moreover, the number of applications for cryptography have increased dramati- cally, as new cryptographic techniques are invented and proven secure. For example, securely transacting with cryptocurrencies such as bitcoin requires modern cryptog- raphy. As another example, hospitals may now share information about patients in a way that protects patient privacy while allowing the hospitals to apply statisti- cal methods assessing the effectiveness of new treatments on the aggregate of the patients. We recommend this book in our MIT class Applied Cryptography. This class is about half undergraduates and half graduate students; past students have said the text was excellent. It will be great to have this new edition available. The approach taken in this text is more pragmatic and engineering-oriented than theory-oriented. It is usable for both classroom use and self-study. This edition of Understanding Cryptography contains much new material; the book has expanded by almost 50% since the first edition. Part of this expansion is due to the expansion of the field (technical), including new problems, and part of the expansion is due to the addition of new references and discussion (historical). Of particular note is the inclusion of new material on “quantum cryptography”: cryptosystems that are specifically designed to resist attacks that are based on the use of quantum computers. Shor’s algorithm (1994) showed that cryptographic al- gorithms that are based on the hardness of factoring the product of two primes, or that are based on the hardness of computing discrete logarithms, are vulnera- ble to polynomial-time attacks using quantum computation. If and when quantum computers become available, cryptographic methods such as RSA or elliptic-curve cryptosystems will become vulnerable. Given the long lead time required to replace cryptosystems in use, planning for a change-over to “quantum-resistant” algorithms has already begun. The (U.S.) National Institute of Standards and Technology vii
viii Foreword textbook covers all three approaches. Indeed, this textbook may be the first to cover PQC (post-quantum cryptography). This textbook also has updated material on “conventional” (non-public-key) cryptography. For example, it includes new and/or updated material on crypto- graphic hash functions (including coverage of SHA-2 and SHA-3), stream ciphers (including Salsa20 and ChaCha), and modes of operation (including authenticated encryption modes). In summary, I recommend this book highly for both undergraduate and graduate classroom use; it can easily be augmented for students with a more theoretical ori- entation. This book is also recommended for self-study, for anyone who wishes to bring themselves up-to-date on where this exciting field is going. December 2023 Ron Rivest has converged on possible standards based on three particular hard problems; this
Preface This is the second edition of Understanding Cryptography. Ever since we released the first edition in 2009, we have been humbled by the many positive responses we received from readers from all over the world. Our goal has always been to make the fascinating but also challenging topic of cryptography accessible and fun to learn. Key concepts of the book are that we focus on cryptography with high practical relevance, and that the necessary mathematical material is accessible for readers with a minimum background in college-level calculus. The fact that Understanding Cryptography has been adopted as textbook by hundreds of universities on all conti- nents (that is, if we ignore Antarctica) and the feedback we received from individual readers and instructors makes us believe that this approach is working. One thing that has changed since the first edition is that it has become abun- dantly clear how important cybersecurity is in our, by now, digital society. Today, seemingly every aspect in our private lives, at work or in governments has become dependent on information technology in one way or another. Even though digital- ization can have many benefits for individuals and society at large, information tech- nology must come with strong security mechanisms in order to prevent malicious manipulations. Here is where cryptography comes into play: It is a key tool for building sound cybersecurity solutions. To this end, cryptographic algorithms have crept into myriads of applications that surround us; examples range from social net- works, smartphones and cloud servers to embedded systems like medical implants, car keys and passports. Emerging applications such as autonomous cars and e-voting will rely even more on strong security mechanisms. Of course, cryptocurrencies and blockchains rely heavily on modern cryptographic algorithms, too. Content Overview The book has many features that make it a unique source for students, practition- ers and researchers. We focus on practical relevance by introducing the majority of cryptographic algorithms that are used in modern real-world applications. With respect to symmetric algorithms, we introduce the block ciphers AES, DES and ix
x Preface triple-DES as well as PRESENT, which is an important example of a lightweight cipher. We also describe three popular stream ciphers. Regarding asymmetric cryp- tography, we cover all three public-key families currently in use: RSA, discrete log- arithm schemes and elliptic curves. In addition, the book introduces hash functions, digital signatures and message authentication codes, or MACs. Beyond core cryp- tographic algorithms, we also discuss topics such as modes of operation, security services and key management. For every cryptographic scheme, up-to-date security estimations and recommendations for key lengths are given. We also discuss the important issue of software and hardware implementation. What’s New The second edition has received major updates and has grown from the 350 pages of the first edition to more than 500 pages. The most noticeable new material is the extensive treatment of post-quantum cryptography, or PQC, in Chapter 12. In the coming years, many applications will need to replace traditional public-key schemes with PQC algorithms. This will be the most comprehensive change in the landscape of cryptography that we have seen in decades. We hope that our introduction to the three most promising PQC families, that is lattice-based, code-based and hash-based schemes, will be helpful in this context. Beside PQC, the 2nd edition also covers the SHA-2 and SHA-3 hash functions, the new stream ciphers Salsa20 and ChaCha, and authenticated encryption. Throughout the book, security parameters and related work have been updated, as well as the Discussion and Further Reading sections that conclude each chapter. The problem sections of all 14 chapters have been extended, too. How to Use the Book The material in this book has evolved over many years and is “classroom proven”. We’ve taught it both as a course for advanced undergraduate students and gradu- ate students in computer science/math/electrical engineering, as well as a first-year undergraduate course for students majoring in our IT security program. We found that one can teach most concepts introduced in the book in a two-semester course, with 90 minutes of lecture time plus 90 minutes of help sessions with exercises per week (total of 10 ECTS credits). In a typical US-style three-credit course, or in a one-semester European course, some of the material should be omitted. Here are some reasonable choices for a one-semester course: Course Curriculum 1 Focus on the application of cryptography, e.g., in an applied course in computer science or a basic course for subsequent security classes, e.g., in a cybersecurity program. A possible curriculum is: Chap. 1; Sects. 2.1–2.2; Chap. 4; Sect. 5.1; Chap. 6; Sects. 7.1–7.3; Sects. 8.1–8.3; Sects. 10.1–10.2; Sects. 11.1–11.3; Sects. 12.1 & 12.4; Sect. 13.1; Sects. 14.1–14.3.
Preface xi Course Curriculum 2 Focus on cryptographic algorithms and their mathematical background, e.g., as a theory course in computer science or a crypto course in a math program. This curriculum also works nicely as preparation for a more theoretical course in cryptography: Chap. 1; Chap. 2; Chap. 4; Chap. 6; Chap. 7; Sects. 8.1 – 8.4; Chap. 9; Sects. 10.1–10.2; Sects. 11.1–11.3; Sects. 12.1, 12.2 & 12.4. More Information There are two online sources related to this book that we can recommend. First, we recorded the two-semester introductory cryptography course that we teach at Ruhr University Bochum (RUB). The main audience for this class are the first- year students of RUB‘s IT Security program, and we tried to make the material as accessible as possible. More than 20 lectures are available on the YouTube channel “Introduction to Cryptography by Christof Paar”: https://www.crypto-textbook.com/video Each lecture takes about 80–90 minutes and closely follows the material in the book. (For the more adventurous reader, there is also a German-language set of videos available in the YouTube channel “Einf¨uhrung in die Kryptographie von Christof Paar”.) Second, we recommend the companion website for the book, containing slide sets for lecturers and solutions to odd-numbers problems of the book: https://www.crypto-textbook.com Trained as engineers, we have worked in applied cryptography and security for more than 20 years and hope that the readers will have as much fun with this fasci- nating field as we’ve had! Christof Paar Jan Pelzl Tim G¨uneysu Bochum, Germany Hamm, Germany Bochum, Germany
Acknowledgements Writing this book would have been impossible without the help of many people. We hope we did not forget anyone in our listing. Help with technical questions was provided by Frederick Armknecht (stream ci- phers), Roberto Avanzi (finite fields and elliptic curves), Eike Kiltz (provable secu- rity), Gregor Leander (block ciphers), Alex May (number theory), Alfred Menezes and Neal Koblitz (history of elliptic curve cryptography), Matt Robshaw (AES) and Damian Weber (discrete logarithms). We are particular grateful to Axel Poschmann who provided the inital section about the PRESENT block cipher. We would also like to thank Conny Robrahn who worked tirelessly on the more than 130 figures in this book. Special thanks for proofreading and the many sugges- tions for improving the material in the second edition go the members of the Embed- ded Security group at the Max Planck Institute for Security and Privacy, Bochum, and the Security Engineering group at Ruhr University Bochum: Nils Albartus, Sven Argo, Steffen Becker, Fabian Buschkowski, Maik Ender, Jakob Feldtkeller, Anna Guinet, Dina Hesse, Simon Klix, Elisabeth Krahmer, Markus Krausz, Georg Land, Johannes Mono, Endres Puschner, Jan Richter-Brockmann, Julian Speith, Paul Staat and Jan Thoma. For the first edition, we are indebted to the members of the Embedded Secu- rity group at Ruhr University Bochum — Andrey Bogdanov, Benedikt Driessen, Thomas Eisenbarth, Stefan Heyse, Markus Kasper, Timo Kasper, Amir Moradi and Daehyun Strobel — who did much of the technical proofreading and provided nu- merous suggestions for improving the presentation of the material. We would like to express our deepest gratitude to Ron Rivest for his willing- ness to provide the foreword. We’d like to thank the people from Springer for their continuous support and encouragement. In particular, thanks to our editors Ronan Nugent and Wayne Wheeler as well as to Michela Castrica. Last but not least we would like to thank all the readers of the first edition who provided valuable feedback regarding improving the text and the problem sets. xiii
Table of Contents 1 Introduction to Cryptography and Data Security . . . . . . . . . . . . . . . . . . 1 1.1 Overview of Cryptology (and This Book) . . . . . . . . . . . . . . . . . . . . . . 2 1.2 Symmetric Cryptography . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 1.2.1 Basics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 1.2.2 Simple Symmetric Encryption: The Substitution Cipher . . . . 7 1.3 Cryptanalysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 1.3.1 General Thoughts on Breaking Cryptosystems . . . . . . . . . . . . 10 1.3.2 How Many Key Bits Are Enough? . . . . . . . . . . . . . . . . . . . . . . 13 1.4 Modular Arithmetic and More Historical Ciphers . . . . . . . . . . . . . . . . 15 1.4.1 Modular Arithmetic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 1.4.2 Integer Rings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 1.4.3 Shift Cipher (or Caesar Cipher) . . . . . . . . . . . . . . . . . . . . . . . . 20 1.4.4 Affine Cipher . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 1.5 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 1.6 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 2 Stream Ciphers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38 2.1.1 Stream Ciphers vs. Block Ciphers . . . . . . . . . . . . . . . . . . . . . . 38 2.1.2 Encryption and Decryption with Stream Ciphers . . . . . . . . . . 40 2.2 Random Numbers and an Unbreakable Stream Cipher . . . . . . . . . . . . 43 2.2.1 Random Number Generators . . . . . . . . . . . . . . . . . . . . . . . . . . . 43 2.2.2 The One-Time Pad . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 2.2.3 Towards Practical Stream Ciphers . . . . . . . . . . . . . . . . . . . . . . 46 2.3 Shift Register-Based Stream Ciphers . . . . . . . . . . . . . . . . . . . . . . . . . . 49 2.3.1 Linear Feedback Shift Registers (LFSRs) . . . . . . . . . . . . . . . . 50 2.3.2 Known-Plaintext Attack Against Single LFSRs . . . . . . . . . . . 53 2.4 Practical Stream Ciphers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55 2.4.1 Salsa20 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55 2.4.2 ChaCha . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59 xv
xvi Table of Contents 2.4.3 Trivium . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61 2.5 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64 2.6 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68 3 The Data Encryption Standard (DES) and Alternatives . . . . . . . . . . . . . 73 3.1 Introduction to DES . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74 3.1.1 Confusion and Diffusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75 3.2 Overview of the DES Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 3.3 Internal Structure of DES . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 3.3.1 Initial and Final Permutation . . . . . . . . . . . . . . . . . . . . . . . . . . . 80 3.3.2 The f Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81 3.3.3 Key Schedule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86 3.4 Decryption . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 88 3.5 Security of DES . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92 3.5.1 Exhaustive Key Search . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93 3.5.2 Analytical Attacks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 3.6 Implementation in Software and Hardware . . . . . . . . . . . . . . . . . . . . . 96 3.7 DES Alternatives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97 3.7.1 The Advanced Encryption Standard (AES) and the AES Finalist Ciphers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97 3.7.2 Triple DES (3DES) and DESX . . . . . . . . . . . . . . . . . . . . . . . . . 98 3.7.3 Lightweight Cipher PRESENT . . . . . . . . . . . . . . . . . . . . . . . . . 99 3.8 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103 3.9 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106 4 The Advanced Encryption Standard (AES) . . . . . . . . . . . . . . . . . . . . . . . 111 4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 112 4.2 Overview of the AES Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 113 4.3 Some Mathematics: A Brief Introduction to Galois Fields . . . . . . . . . 114 4.3.1 Existence of Finite Fields . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 116 4.3.2 Prime Fields . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117 4.3.3 Extension Fields GF(2m) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119 4.3.4 Addition and Subtraction in GF(2m) . . . . . . . . . . . . . . . . . . . . 120 4.3.5 Multiplication in GF(2m) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120 4.3.6 Inversion in GF(2m) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123 4.4 Internal Structure of AES . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124 4.4.1 Byte Substitution Layer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 4.4.2 Diffusion Layer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128 4.4.3 Key Addition Layer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130 4.4.4 Key Schedule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131 4.5 Decryption . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135 4.6 Implementation in Software and Hardware . . . . . . . . . . . . . . . . . . . . . 140 4.7 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141
Table of Contents xvii 4.8 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 142 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 143 5 More About Block Ciphers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147 5.1 Modes of Operation for Encryption and Authentication . . . . . . . . . . . 148 5.1.1 Electronic Codebook Mode (ECB) . . . . . . . . . . . . . . . . . . . . . . 149 5.1.2 Cipher Block Chaining Mode (CBC) and Initialization Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153 5.1.3 Output Feedback Mode (OFB) . . . . . . . . . . . . . . . . . . . . . . . . . 155 5.1.4 Cipher Feedback Mode (CFB) . . . . . . . . . . . . . . . . . . . . . . . . . 156 5.1.5 Counter Mode (CTR) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157 5.1.6 XTS-AES . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159 5.2 Exhaustive Key Search Revisited . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 161 5.3 Increasing the Security of Block Ciphers . . . . . . . . . . . . . . . . . . . . . . . 162 5.3.1 Double Encryption and Meet-in-the-Middle Attack . . . . . . . . 163 5.3.2 Triple Encryption . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 165 5.3.3 Key Whitening . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167 5.4 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168 5.5 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 170 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 171 6 Introduction to Public-Key Cryptography . . . . . . . . . . . . . . . . . . . . . . . . 177 6.1 Symmetric vs. Asymmetric Cryptography . . . . . . . . . . . . . . . . . . . . . . 178 6.2 Practical Aspects of Public-Key Cryptography . . . . . . . . . . . . . . . . . . 182 6.2.1 Security Mechanisms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 183 6.2.2 The Remaining Problem: Authenticity of Public Keys . . . . . 184 6.2.3 Important Public-Key Algorithms . . . . . . . . . . . . . . . . . . . . . . 184 6.2.4 Key Lengths and Security Levels . . . . . . . . . . . . . . . . . . . . . . . 185 6.3 Essential Number Theory for Public-Key Algorithms . . . . . . . . . . . . 186 6.3.1 Euclidean Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 187 6.3.2 Extended Euclidean Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . 189 6.3.3 Euler’s Phi Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195 6.3.4 Fermat’s Little Theorem and Euler’s Theorem . . . . . . . . . . . . 197 6.4 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 199 6.5 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 200 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201 7 The RSA Cryptosystem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205 7.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 206 7.2 Encryption and Decryption . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 206 7.3 Key Generation and Proof of Correctness . . . . . . . . . . . . . . . . . . . . . . 207 7.4 Encryption and Decryption: Fast Exponentiation . . . . . . . . . . . . . . . . 211 7.5 Speed-Up Techniques for RSA . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 215 7.5.1 Fast Encryption with Short Public Exponents . . . . . . . . . . . . . 215 7.5.2 Fast Decryption with the Chinese Remainder Theorem . . . . . 216
xviii Table of Contents 7.6 Finding Large Primes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 219 7.6.1 How Common Are Primes? . . . . . . . . . . . . . . . . . . . . . . . . . . . . 220 7.6.2 Primality Tests . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 221 7.7 RSA in Practice: Padding . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 224 7.8 Key Encapsulation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 226 7.9 Attacks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 228 7.10 Implementation in Software and Hardware . . . . . . . . . . . . . . . . . . . . . 230 7.11 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 232 7.12 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 234 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 235 8 Cryptosystems Based on the Discrete Logarithm Problem . . . . . . . . . . 241 8.1 Diffie–Hellman Key Exchange . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 242 8.2 Some Abstract Algebra . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 244 8.2.1 Groups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 244 8.2.2 Cyclic Groups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 246 8.2.3 Subgroups . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 250 8.3 The Discrete Logarithm Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 252 8.3.1 The Discrete Logarithm Problem in Prime Fields . . . . . . . . . . 252 8.3.2 The Generalized Discrete Logarithm Problem . . . . . . . . . . . . 253 8.3.3 Attacks Against the Discrete Logarithm Problem . . . . . . . . . . 255 8.4 Security of the Diffie–Hellman Key Exchange . . . . . . . . . . . . . . . . . . 260 8.5 The Elgamal Encryption Scheme . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 261 8.5.1 From Diffie–Hellman Key Exchange to Elgamal Encryption 261 8.5.2 The Elgamal Protocol . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 262 8.5.3 Computational Aspects . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 264 8.5.4 Security . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 265 8.6 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 267 8.7 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 269 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 270 9 Elliptic Curve Cryptosystems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 277 9.1 How to Compute with Elliptic Curves . . . . . . . . . . . . . . . . . . . . . . . . . 278 9.1.1 Definition of Elliptic Curves . . . . . . . . . . . . . . . . . . . . . . . . . . . 279 9.1.2 Group Operations on Elliptic Curves . . . . . . . . . . . . . . . . . . . . 281 9.2 Building a Discrete Logarithm Problem with Elliptic Curves . . . . . . 285 9.3 Diffie–Hellman Key Exchange with Elliptic Curves . . . . . . . . . . . . . . 289 9.4 Security . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 291 9.5 Implementation in Software and Hardware . . . . . . . . . . . . . . . . . . . . . 292 9.6 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 293 9.7 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 295 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 296
Table of Contents xix 10 Digital Signatures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 299 10.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 300 10.1.1 Odd Colors for Cars, or: Why Symmetric Cryptography Is Not Sufficient . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 300 10.1.2 Principles of Digital Signatures . . . . . . . . . . . . . . . . . . . . . . . . 301 10.1.3 Security Services . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 303 10.1.4 Applications of Digital Signatures . . . . . . . . . . . . . . . . . . . . . . 305 10.2 The RSA Signature Scheme . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 306 10.2.1 Schoolbook RSA Digital Signature . . . . . . . . . . . . . . . . . . . . . 306 10.2.2 Computational Aspects . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 308 10.2.3 Security . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 309 10.3 The Elgamal Digital Signature Scheme . . . . . . . . . . . . . . . . . . . . . . . . 312 10.3.1 Schoolbook Elgamal Digital Signature . . . . . . . . . . . . . . . . . . 312 10.3.2 Computational Aspects . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 315 10.3.3 Security . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 315 10.4 The Digital Signature Algorithm (DSA) . . . . . . . . . . . . . . . . . . . . . . . . 318 10.4.1 The DSA Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 318 10.4.2 Computational Aspects . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 322 10.4.3 Security . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 323 10.5 The Elliptic Curve Digital Signature Algorithm (ECDSA) . . . . . . . . 324 10.5.1 The ECDSA Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324 10.5.2 Computational Aspects . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 327 10.5.3 Security . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 328 10.6 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 329 10.7 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 330 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 331 11 Hash Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 335 11.1 Motivation: Signing Long Messages . . . . . . . . . . . . . . . . . . . . . . . . . . . 336 11.2 Security Requirements of Hash Functions . . . . . . . . . . . . . . . . . . . . . . 339 11.2.1 Preimage Resistance or One-Wayness . . . . . . . . . . . . . . . . . . . 339 11.2.2 Second Preimage Resistance or Weak Collision Resistance . 340 11.2.3 Collision Resistance and the Birthday Attack . . . . . . . . . . . . . 341 11.3 Overview of Hash Algorithms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 346 11.3.1 Hash Functions from Block Ciphers . . . . . . . . . . . . . . . . . . . . 347 11.3.2 The Dedicated Hash Functions SHA-1, SHA-2 and SHA-3 . 349 11.4 The Secure Hash Algorithm SHA-2 . . . . . . . . . . . . . . . . . . . . . . . . . . . 351 11.4.1 SHA-256 Preprocessing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 352 11.4.2 The SHA-256 Compression Function . . . . . . . . . . . . . . . . . . . 353 11.4.3 Implementation in Software and Hardware . . . . . . . . . . . . . . . 356 11.5 The Secure Hash Algorithm SHA-3 . . . . . . . . . . . . . . . . . . . . . . . . . . . 357 11.5.1 High-Level View of SHA-3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . 358 11.5.2 Suffix, Padding and Output Generation . . . . . . . . . . . . . . . . . . 360 11.5.3 The Function Keccak- f (or the Keccak- f Permutation) . . . . 361 11.5.4 Other Cryptographic Functions Based on Keccak . . . . . . . . . 367
xx Table of Contents 11.5.5 Implementation in Software and Hardware . . . . . . . . . . . . . . . 368 11.6 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 369 11.7 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 374 12 Post-Quantum Cryptography . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 379 12.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 380 12.1.1 Quantum Computing and Cryptography . . . . . . . . . . . . . . . . . 380 12.1.2 Quantum-Secure Asymmetric Cryptosystems . . . . . . . . . . . . . 383 12.1.3 The Use of Uncertainty in Cryptography . . . . . . . . . . . . . . . . . 384 12.2 Lattice-Based Cryptography . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 386 12.2.1 The Learning With Errors (LWE) Problem . . . . . . . . . . . . . . . 389 12.2.2 A Simple LWE-Based Encryption System . . . . . . . . . . . . . . . 391 12.2.3 The Ring Learning With Errors Problem . . . . . . . . . . . . . . . . . 399 12.2.4 Ring-LWE Encryption Scheme . . . . . . . . . . . . . . . . . . . . . . . . . 401 12.2.5 LWE in Practice . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 406 12.2.6 Final Remarks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 409 12.3 Code-Based Cryptography . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 410 12.3.1 Linear Codes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 411 12.3.2 The Syndrome Decoding Problem . . . . . . . . . . . . . . . . . . . . . . 417 12.3.3 Encryption Schemes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 419 12.3.4 Suitable Choices of Codes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 427 12.3.5 Final Remarks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 429 12.4 Hash-Based Cryptography . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 430 12.4.1 One-Time Signatures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 430 12.4.2 Many-Time Signatures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 443 12.4.3 Final Remarks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 452 12.5 PQC Standardization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 453 12.6 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 454 12.7 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 457 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 458 13 Message Authentication Codes (MACs) . . . . . . . . . . . . . . . . . . . . . . . . . . . 465 13.1 Principles of Message Authentication Codes . . . . . . . . . . . . . . . . . . . . 466 13.2 MACs from Hash Functions: HMAC . . . . . . . . . . . . . . . . . . . . . . . . . . 468 13.3 MACs from Block Ciphers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 472 13.3.1 CBC-MAC . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 472 13.3.2 Cipher-based MAC (CMAC) . . . . . . . . . . . . . . . . . . . . . . . . . . 473 13.3.3 Authenticated Encryption: The Counter with Cipher Block Chaining-Message Authentication Code (CCM) . . . . . . . . . . 474 13.3.4 Authenticated Encryption: The Galois Counter Mode (GCM)476 13.3.5 Galois Counter Message Authentication Code (GMAC) . . . . 478 13.4 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 478 13.5 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 479 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 480
Table of Contents xxi 14 Key Management . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 483 14.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 484 14.2 Key Derivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 486 14.3 Key Establishment Using Symmetric-Key Techniques . . . . . . . . . . . . 490 14.3.1 Key Establishment with a Key Distribution Center . . . . . . . . 491 14.3.2 Needham-Schroeder Protocol . . . . . . . . . . . . . . . . . . . . . . . . . . 495 14.3.3 Remaining Problems with Symmetric-Key Distribution . . . . 496 14.4 Key Establishment Using Asymmetric Techniques . . . . . . . . . . . . . . . 497 14.4.1 Man-in-the-Middle Attack . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 498 14.4.2 Certificates . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 500 14.5 Public-Key Infrastructures (PKIs) and CAs . . . . . . . . . . . . . . . . . . . . . 504 14.5.1 Certificate Chains . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 505 14.5.2 Certificate Revocation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 506 14.6 Practical Aspects of Key Management . . . . . . . . . . . . . . . . . . . . . . . . . 509 14.7 Discussion and Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 511 14.8 Lessons Learned . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 514 Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 515 References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 521 Index . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 535
Chapter 1 Introduction to Cryptography and Data Security This section will introduce the most important terms of modern cryptology and will teach an important lesson about proprietary vs. openly known algorithms. We will also introduce modular arithmetic, which is useful for historical ciphers and of major importance in modern public-key cryptography. In this chapter you will learn: The general rules of cryptography Key lengths for short-, medium- and long-term security The different ways of attacking ciphers A few historical ciphers and on the way we will learn about modular arithmetic Why one should only use well-established cryptographic algorithms 1 C. Paar et al., Understanding Cryptography, https://doi.org/10.1007/978-3-662-69007-9_1 © The Editor(s) (if applicable) and The Author(s), under exclusive license to Springer-Verlag GmbH, DE, part of Springer Nature 2024
2 1 Introduction to Cryptography and Data Security 1.1 Overview of Cryptology (and This Book) The book at hand provides an introduction to cryptography. This is part of the broader area of cybersecurity, which deals with the protection of digital informa- tion against misuse. Even though cybersecurity is a complex field that encompasses technical aspects as well as organizational and human ones, almost all IT security solutions in practice employ cryptography as a crucial module. A rough analogy comes from the automotive domain: If cybersecurity is a car, cryptography is the en- gine. Even though there are obviously many parts and technologies that are needed for a car, every automobile relies on an engine as a central component. The same holds in the security domain: It is hard to build secure digital systems without cryp- tographic algorithms. As we know from almost daily reports about successful hacks against IT systems, cybersecurity is difficult to achieve. In this context it is impor- tant to bear in mind that today’s cryptography is usually the most secure part of a cybersecurity solution. This book is primarily concerned with modern cryptographic algorithms, also referred to as cryptographic primitives or ciphers. If we hear the word cryptography our first associations might be cryptocurren- cies, end-to-end encryption for the instant messenger running on our smartphone or secure website access. Perhaps we go back a little bit in history and think about the famous attack against the German Enigma encryption machine during World War II (Figure 1.1). In any case, cryptography seems closely linked to modern elec- Fig. 1.1 The German Enigma encryption machine (reproduced with permission of the Deutsches Museum, Munich) tronic communication. However, cryptography is a rather old business, with early examples dating back to about 2000 B.C., when non-standard “secret” hieroglyphics were used in ancient Egypt. Since Egyptian times cryptography has been used in one form or another in many, if not most, cultures that developed written language. For
1.1 Overview of Cryptology (and This Book) 3 instance, there are documented cases of secret writing in ancient Greece, namely the scytale of Sparta (Figure 1.2), or the famous Caesar cipher in ancient Rome, about which we will learn later in this chapter. This book, however, strongly focuses on Fig. 1.2 Scytale of Sparta modern cryptographic methods and also teaches many data security issues and their relationship with cryptography. Let’s now have a look at the field of cryptography, shown in Figure 1.3. The first Fig. 1.3 Overview of the field of cryptology thing that we notice is that the most general term is cryptology and not cryptography. Cryptology splits into two main branches: Cryptography is the science of securing communication against an adversary. Historically, the main goal of crypography was to hide the meaning of a message. Today, however, cryptography is also used for many other security goals such as the integrity and authenticity of messages. Cryptanalysis is the science and sometimes art of breaking cryptosystems. You might think that code breaking is for the intelligence community or perhaps or- ganized crime, and should not be included in a serious classification of a sci- entific discipline. However, most cryptanalysis nowadays is done by respectable
4 1 Introduction to Cryptography and Data Security researchers in academia. Cryptanalysis is of central importance for modern cryp- tosystems: Without people who try to break our cryptographic methods, we will never know whether they are really secure or not. This issue is discussed in more detail in Section 1.3. Because cryptanalysis is the only way to ensure that a cryptosystem is secure, it is an integral part of cryptology. Nevertheless, the focus of this book is on cryptography: We introduce the most important practical cryptographic algorithms in detail. These are all ciphers that have withstood cryptanalysis for a long time, in most cases for several decades. In the case of cryptanalysis we will mainly restrict ourselves to providing state-of-the-art results with respect to breaking the crypto- graphic algorithms that are introduced, e.g., the factoring record for breaking the RSA scheme. Let’s now go back to Figure 1.3. Cryptography itself splits into three main branches: Symmetric Algorithms are what many people assume cryptography is about: Two parties have an encryption and decryption method for which they share a secret key. All cryptography from ancient times until 1976 was exclusively based on symmetric methods. Symmetric ciphers are still in widespread use, especially for actual data encryption and integrity checking of messages. Asymmetric (or Public-Key) Algorithms In 1976 an entirely different type of cipher was introduced by Whitfield Diffie, Martin Hellman and Ralph Merkle. In public-key cryptography, two keys exist: A user possesses a secret key as in symmetric cryptography but also a public key. Asymmetric algorithms can be used for applications such as digital signatures and key establishment but also for classical data encryption. Cryptographic Protocols Roughly speaking, cryptographic protocols realize more complex security functions through the use of cryptographic algorithms. Symmetric and asymmetric algorithms can be viewed as building blocks with which applications such as secure internet communication can be realized. The Transport Layer Security (TLS) scheme, which is used in every web browser, is an example of a cryptographic protocol. Strictly speaking, hash functions, which will be introduced in Chapter 11, form a third class of algorithms but at the same time they share many properties with symmetric ciphers. In the majority of cryptographic applications in practical systems, symmetric and asymmetric algorithms (and often also hash functions) are all used together. These are sometimes referred to as hybrid schemes. The reason for using both families of algorithms is that each has specific strengths and weaknesses. The main focus of this book is on symmetric and asymmetric algorithms, as well as hash functions. However, we will also introduce basic security protocols. In par- ticular, we will introduce several key establishment protocols and discuss what can be achieved with cryptographic protocols, including confidentiality of data, integrity of data, authentication of data, user identification, etc.
1.2 Symmetric Cryptography 5 1.2 Symmetric Cryptography This section deals with the concept of symmetric ciphers and introduces the historic substitution cipher. Using the substitution cipher as an example, we will learn the difference between brute-force and analytical attacks. 1.2.1 Basics Symmetric cryptographic schemes are also referred to as symmetric-key, secret-key and single-key schemes or algorithms. Symmetric cryptography is best introduced with an easy-to-understand problem: There are two users, Alice and Bob, who want to communicate over an insecure channel (Figure 1.4). The term channel might sound a bit abstract but it is just a general term for the communication link: This can be the internet, a stretch of air in the case of smartphones or a Wi-Fi home network, or any other communication media you can think of. The actual problem starts with the bad guy, Oscar1, who has access to the channel, for instance, by hacking into an internet router or by listening to the radio signals of a Wi-Fi communication. This type of unauthorized listening is called eavesdropping. Obviously, there are many situations in which Alice and Bob would prefer to communicate without Oscar listening. For instance, if Alice and Bob represent the headquarters and the research office of a pharmaceutical company, and they are transmitting documents containing their strategy for the development of a new pharmaceutical drug over the next few years, these documents should not get into the hands of competitors, or of foreign intelligence agencies for that matter. Fig. 1.4 Communication over an insecure channel In this situation, symmetric cryptography offers a powerful solution: Alice en- crypts her message x using a symmetric algorithm, yielding the ciphertext y. Bob receives the ciphertext and decrypts the message. Decryption is, thus, the inverse 1 The name Oscar was chosen to remind us of the word opponent.
6 1 Introduction to Cryptography and Data Security process of encryption (Figure 1.5). What is the advantage? If we have a strong en- cryption algorithm, the ciphertext will look like random bits and Oscar will not be able to obtain any useful information from it. Fig. 1.5 Symmetric-key cryptosystem The variables x, y and k in Figure 1.5 have special names in cryptography: x is called the plaintext or cleartext, y is called the ciphertext, k is called the key. Remark that the set of all possible keys is called the key space. The system needs a secure channel for distribution of the key between Alice and Bob. The secure channel shown in Figure 1.5 can, for instance, be a human who is transporting the key in a wallet between Alice and Bob. This is, of course, a cumbersome method. An example where this method works nicely is the pre-shared keys used in Wi-Fi Protected Access (WPA) encryption in wireless LANs. Later in this book we will learn methods for establishing keys over insecure channels. In any case, the key has only to be transmitted once between Alice and Bob and can then be used for securing many subsequent communications. One important and also counterintuitive fact in this situation is that both the en- cryption and the decryption algorithms are publicly known. It seems that keeping the encryption algorithm secret should make the whole system harder to break. How- ever, secret algorithms also mean less intensively tested algorithms: The only way to find out whether an encryption method is strong, i.e., cannot be broken by a de- termined attacker, is to make it public and have it analyzed by other cryptographers. Please see Section 1.3 for more discussion on this topic. The only thing that should be kept secret in a sound cryptosystem is the key. Remarks: 1. It seems very likely that most modern cryptographic algorithms can not be broken by anybody on planet Earth, including big intelligence agencies. This assumes,
1.2 Symmetric Cryptography 7 however, that the algorithm is used correctly. Especially, we have to ensure that an attacker does not get hold of the key. Of course, once Oscar knows the key, he can easily decrypt the message since the algorithm is publicly known. Hence it is crucial to note that the problem of transmitting a message se- curely is reduced to the problems of transmitting a key secretly and of stor- ing the key in a secure fashion. 2. In this scenario we only consider the problem of confidentiality, that is, of hiding the contents of the message from an eavesdropper. We will see later in this book that there are many other things we can do with cryptography, such as preventing Oscar from making unnoticed changes to the message (message integrity) or ensuring that a message really comes from Alice (sender authentication). 1.2.2 Simple Symmetric Encryption: The Substitution Cipher We will now learn one of the simplest methods for encrypting text, the substitution (= replacement) cipher. Historically this type of cipher has been widely used, and it is a good illustration of basic cryptography. We will use the substitution cipher for learning some important facts about key lengths and about different ways of attacking cryptographic algorithms. The goal of the substitution cipher is the encryption of text (as opposed to bits in modern digital systems). The idea is very simple: We substitute each letter of the alphabet with another one. Example 1.1. Plaintext Ciphertext A → k B → d C → w · · · For instance, the pop group ABBA would be encrypted as kddk. We assume that we choose the substitution table completely randomly, so that an attacker is not able to guess it. Note that the substitution table is the key of this cryptosystem. As always in symmetric cryptography, the key, i.e., the substitution table, has to be distributed between Alice and Bob in a secure fashion. Example 1.2. Let’s look at another ciphertext: iq ifcc vqqr fb rdq vfllcq na rdq cfjwhwz hr bnnb hcc hwwhbsqvqbre hwq vhlq
8 1 Introduction to Cryptography and Data Security This does not seem to make too much sense and looks like decent cryptography. However, the substitution cipher is not secure at all! Let’s look at ways of breaking the cipher. First Attack: Brute-Force Attack or Exhaustive Key Search Brute-force attacks treat the cipher as a black box. They are based on a simple con- cept: Oscar, the attacker, has the ciphertext from eavesdropping on the channel and happens to have a short piece of plaintext, e.g., the header of a file that was en- crypted. Oscar now simply decrypts the first piece of ciphertext with all possible keys. Again, the key for this cipher is the substitution table. If the resulting plaintext matches the short piece of plaintext, he knows that he has found the correct key. Definition 1.2.1 Basic Exhaustive Key Search or Brute-Force At- tack Let (x, y) denote the pair of plaintext and ciphertext, and let K = {k1, ..., kκ } be the key space of all possible keys ki. A brute-force attack now checks for every ki ∈ K whether dki (y) ? = x. If the equality holds, a possible correct key is found; if not, proceed with the next key. In practice, a brute-force attack can be more complicated because incorrect keys can give false positive results. We will address this issue in Section 5.2. It is important to note that a brute-force attack against symmetric ciphers is al- ways possible in principle. Whether it is feasible in practice depends on the key space, i.e., on the number of possible keys that exist for a given cipher. If testing all the keys on many modern computers takes too much time, i.e., hundreds or thou- sands of years, the cipher is computationally secure against a brute-force attack. More on computational security will be said in Section 2.2.3. Let’s determine the key space of the substitution cipher: When choosing the re- placement for the first letter A, we randomly choose one letter from the 26 letters of the alphabet (in the example above we chose k). The replacement for the next al- phabet letter B was randomly chosen from the remaining 25 letters, etc. Thus there exist the following number of different substitution tables: key space of the substitution cipher = 26 · 25 · · · 3 · 2 · 1 = 26! ≈ 288 That means the key space has roughly a size of 288, which is equal to the key space of a cipher that has a key consisting of 88 bits. Even with hundreds of thousands of high-end PCs such a search would take several decades! Thus, we are tempted to
1.2 Symmetric Cryptography 9 conclude that the substitution cipher is secure. But this is incorrect because there is another, more powerful, attack, which will be described in the following. Second Attack: Letter Frequency Analysis First we note that the brute-force attack from above treats the cipher as a black box, i.e., we do not analyze the internal structure of the cipher. The substitution cipher can easily be broken by such an analytical attack. The major weakness of the cipher is that each plaintext symbol always maps to the same ciphertext symbol. That means that the statistical properties of the plaintext are preserved in the ciphertext. If we go back to the second example we observe that the letter q occurs most frequently in the text. From this we know that q must be the substitution for one of the frequent letters in the English language. For practical attacks, the following properties of language can be exploited: 1. Determine the frequency of every ciphertext letter. The frequency distribution, usually quite stable even for relatively short pieces of encrypted text, will be close to that of the given language in general. In particular, the most frequent letters can often easily be spotted in ciphertexts. For instance, in English E is the most frequent letter (about 13%), T is the second most frequent letter (about 9%), A is the third most frequent letter (about 8%), and so on. Table 1.1 lists the letter frequency distribution of English. Table 1.1 Relative letter frequencies of the English language Letter Frequency Letter Frequency A 0.0817 N 0.0675 B 0.0150 O 0.0751 C 0.0278 P 0.0193 D 0.0425 Q 0.0010 E 0.1270 R 0.0599 F 0.0223 S 0.0633 G 0.0202 T 0.0906 H 0.0609 U 0.0276 I 0.0697 V 0.0098 J 0.0015 W 0.0236 K 0.0077 X 0.0015 L 0.0403 Y 0.0197 M 0.0241 Z 0.0007 2. The method above can be generalized by looking at pairs or triples, or quadru- ples, and so on of ciphertext symbols. For instance, in English (and some other European languages), the letter Q is almost always followed by a U. This behavior can be exploited to detect the substitution of the letter Q and the letter U. 3. If we assume that word separators, which means “blanks”, have been found (which is only sometimes the case), one can often detect frequent short words
10 1 Introduction to Cryptography and Data Security such as THE, AND, etc. Once we have identified one of these words, we imme- diately know three letters (or whatever the length of the word is) for the entire text. In practice, the three techniques listed above are often combined to break substi- tution ciphers. Example 1.3. If we analyze the encrypted text from Example 1.2, we obtain: WE WILL MEET IN THE MIDDLE OF THE LIBRARY AT NOON ALL ARRANGEMENTS ARE MADE Lesson learned Good ciphers should hide the statistical properties of the encrypted plaintext. The ciphertext symbols should appear to be random. Also, a large key space alone is not sufficient for a strong encryption function. 1.3 Cryptanalysis This section deals with recommended key lengths of symmetric ciphers and differ- ent ways of attacking cryptographic algorithms. It is stressed that a cipher should be secure even if the attacker knows the details of the algorithm. 1.3.1 General Thoughts on Breaking Cryptosystems If we ask someone with some technical background what breaking ciphers is about, he/she will most likely say that code breaking has to do with heavy mathematics, smart people and large computers. We have images in mind of the British code breakers during World War II, attacking the German Enigma cipher with extremely smart mathematicians (the famous computer scientist Alan Turing headed the ef- forts) and room-sized electro-mechanical computers. However, in practice there are also other methods for code breaking. For a secure cryptosystem, it is important (1) to use sound cryptographic algorithms and protocols and (2) to use correct imple- mentations of the algorithms. Let’s look at different ways of breaking cryptosystems in the real world shown in Figure 1.6. Classical Cryptanalysis Classical cryptanalysis attempts to break a cipher by analyzing the inputs and out- puts. We recall from the earlier discussion that cryptanalysis can be divided into analytical attacks, which exploit the internal structure of the encryption method,
1.3 Cryptanalysis 11 Fig. 1.6 Overview of cryptanalysis and brute-force attacks, which treat the encryption algorithm as a black box and test all possible keys. The specific goal of the adversary can vary but in most cases Oscar attempts to recover the plaintext x from the ciphertext y or he attempts to recover the key k from the ciphertext y. Especially for analytical attacks it is helpful to look at what information the opponent has in addition to the ciphertext. The main classes of attacks are: Ciphertext-only attack: The adversary has only access to the ciphertext. Known-plaintext attack: In addition to the ciphertext, the adversary also knows some pieces of the plaintext (e.g., header information of an encrypted file or email). Chosen-plaintext attack: The adversary can choose the plaintext that is being en- crypted and also has access to the corresponding ciphertext. This can for instance be the case when he has access to a decryption device such as a smart card and he attempts to recover the secret key. Chosen-ciphertext attack: The adversary can choose ciphertexts and also obtains the corresponding plaintexts. Again, the goal is typically to recover the secret key. This list is not exhaustive; additional attacks include adaptive chosen-plaintext and adaptive chosen-ciphertext attacks or the related-key attack. Implementation Attacks Side-channel analysis can be used to extract a secret key by observing the behavior of a cryptographic implementation, e.g., an integrated circuit or a piece of software. One family of attacks uses the electrical power consumption or electromagnetic ra- diation of the CPU that computes the cryptographic algorithms as sidechannels. The attacker records the power or electromagnetic traces and applies signal processing techniques to recover the key. Related attacks are based on timing side-channels, in which the adversary measures the run time behavior of a cryptographic implemen-
12 1 Introduction to Cryptography and Data Security tation and attempts to compute the key from the timing measurements. All of these attacks are mainly used against devices to which an attacker has physical access such as smart cards, smartphones or IoT devices.2 Another family of attacks exploits software side-channels. They are primarily relevant if different processes are running on a computer, e.g., in cloud computing. The assumption is that the adversary controls one process with which he is able to learn secret values such as cryptographic keys from another process. To gain infor- mation, the hostile process exploits effects such as timing behavior or cache access patterns. A main mechanism for preventing software side-channels is to ensure that cryptographic implementations have a constant run time, independent of any secret value. Social Engineering Attacks Bribing, blackmailing, tricking or classical espionage can be used to obtain a secret key by involving humans. For instance, forcing someone to reveal his/her secret key, e.g., by holding a gun to his/her head, can be quite successful. Another, less violent, attack is to simply call the victim by phone and say: “This is the IT department of your company. For important software updates we need your password”. It is always surprising how many people are na¨ıve enough to actually give out their passwords in such situations. Even though both implementation attacks and social engineering attacks can be quite powerful in practice, this book mainly assumes attacks based on mathematical cryptanalysis and brute-force attacks. We note that the list of attacks against cryptographic systems is certainly not exhaustive. For instance, malware on a computer can also reveal secret keys in software systems. You might think that many of these attacks, especially social engineering and implementation attacks, are “unfair” but there is little fairness in real-world cryptography. If people want to break your IT system, they are already breaking the rules and are, thus, unfair. The major point to learn here is: An attacker always looks for the weakest link in your cryptosystem. That means we have to choose strong algorithms and we have to make sure that all other attacks such as social engineering and implementation attacks are not feasible. Solid cryptosystems should adhere to Kerckhoffs’ Principle, postulated by Au- guste Kerckhoffs in 1883. 2 Note that most modern hardware tokens that are security sensitive, such as smart cards used for payment, have built-in countermeasures against sidechannel attacks and are very hard to break.
1.3 Cryptanalysis 13 Definition 1.3.1 Kerckhoffs’ Principle A cryptosystem should be secure even if the attacker (Oscar) knows all details about the system, with the exception of the secret key. In particular, the system should be secure when the attacker knows the encryption and decryption algorithms. Some background information on the principle can be found in the Further Reading, Section 1.5. Important Remark: Kerckhoffs’ Principle is counterintuitive! It is extremely tempting to design a system that appears to be more secure because we keep the de- tails hidden. This is called security by obscurity. However, experience and military history has shown over time that such systems are almost always weak, and they are very often broken easily as soon as the secret design has been reverse-engineered or leaked out through other means. An instructive case study for this is the attack on Mifare chipcards. This type of chipcard had been used millionfold in applications for contactless payment, e.g., in the original Oyster card used for London’s public transportation system. Its security was based on a cipher which was kept secret. This worked “well” for several years. However, after reverse-engineering the cipher, re- searchers quickly found several ways of attacking the algorithm, both with classical cryptanalysis and implementation attacks. This lead to severe security problems for the real-world systems that were based on Mifare. For this reason, cryptographic algorithms must provide security even if an attacker gets to known to all internal details except for the key. 1.3.2 How Many Key Bits Are Enough? During the 1990s there was much public discussion about the key length of ciphers. Before we provide some guidelines, there are two crucial aspects to remember: 1. The discussion of key lengths for symmetric cryptographic algorithms is only rel- evant if a brute-force attack is the best known attack. As we saw in Section 1.2.2 during the security analysis of the substitution cipher, if there is an analytical attack that works, a large key space does not help at all. Of course, if there is the possibility of social engineering or implementation attacks, a long key also does not help. 2. The key lengths for symmetric and asymmetric algorithms are dramatically dif- ferent. For instance, a 128-bit symmetric key provides roughly the same security as a 3072-bit RSA (RSA is a popular asymmetric algorithm) key. Both facts are often misunderstood, especially in the semitechnical literature. Table 1.2 gives a rough indication of the security of symmetric ciphers with re- spect to brute-force attacks. As described in Section 1.2.2, a large key space is a nec- essary but not sufficient condition for a secure symmetric cipher. The cipher must
14 1 Introduction to Cryptography and Data Security also be strong against analytical attacks. The table mentions quantum computers. Table 1.2 Estimated time for successful brute-force attacks on symmetric cipher with different key lengths Key length Security estimation 56–64 bits short term: a few hours or days 112–128 bits long term: several decades in the absence of quantum computers 256 bits long term: several decades, even with quantum computers that run the currently known quantum computing algorithms The role that they play for the cryptanalysis of symmetric ciphers is discussed in Section 12.1.1. Foretelling the Future Of course, predicting the future tends to be tricky: We can- not really foresee new technical or theoretical developments with certainty. As you can imagine, it is very hard to know what kinds of computers will be available in the year 2050. For medium-term predictions, Moore’s Law is often assumed. Roughly speaking, Moore’s Law states that computing power doubles every 18 months3 while the costs stay constant. This has the following implications in cryptography: If today we need one month and computers worth $1,000,000 to break a cipher X, then: The cost for breaking the cipher will be $500,000 in 18 months (since we only have to buy half as many computers), $250,000 in 3 years, $125,000 in 4.5 years, and so on. It is important to stress that Moore’s law is an exponential function. In 15 years, i.e., after 10 iterations of computer power doubling, we can do 210 = 1024 times as many computations for the same money we would need to spend today. Stated differently, we only need to spend about 1/1000th of today’s money to do the same computation. In the example above that means that we can break cipher X in 15 years within one month at a cost of about $1, 000, 000/1024 ≈ $1000. Alternatively, with $1,000,000, an attack can be accomplished within 45 minutes in 15 years from now. Moore’s law behaves similarly to a bank account which pays a 100% interest rate every 18 months: The compound interest grows very, very quickly. Unfortu- nately, there are few trustworthy banks which offer such an interest rate. 3 In the literature, the doubling period of Moore’s law is sometimes alternatively given as 24 months. In the security context, it barely matters what exactly the doubling period is. The crucial fact is that computing power grows exponentially over time.
1.4 Modular Arithmetic and More Historical Ciphers 15 1.4 Modular Arithmetic and More Historical Ciphers In this section we use two historical ciphers to introduce modular arithmetic with integers. Even though the historical ciphers are no longer relevant, modular arith- metic is extremely important in modern cryptography, especially for asymmetric algorithms. Ancient ciphers date back to Egypt, where substitution ciphers were used. A very popular special case of the substitution cipher is the Caesar cipher, which is said to have been used by Julius Caesar to communicate with his army. The Caesar cipher simply shifts the letters in the alphabet by a constant number of steps. When the end of the alphabet is reached, the letters repeat in a cyclic way, similarly to numbers in modular arithmetic. To make computations with letters more practicable, we can assign each letter of the alphabet a number. By doing so, an encryption with the Caesar cipher simply becomes a (modular) addition with a fixed value. Instead of just adding constants, a multiplication with a constant can be applied as well. This leads us to the affine cipher. Both the Caesar cipher and the affine cipher will now be discussed in more detail. 1.4.1 Modular Arithmetic Almost all cryptographic algorithms, both symmetric ciphers and asymmetric ci- phers, are based on arithmetic within a finite number of elements. Most number sets we are used to, such as the set of natural numbers or the set of real numbers, are infinite. In the following we introduce modular arithmetic, which is a simple way of performing arithmetic on a finite set of integers. Let’s look at an example of a finite set of integers from everyday life: Example 1.4. Consider the hours on a clock. If you keep adding one hour, you ob- tain: 1h, 2h, 3h, . . . , 11h, 12h, 1h, 2h, 3h, . . . , 11h, 12h, 1h, 2h, 3h, . . . Even though we keep adding one hour, we never leave the set. Let’s look at a general way of dealing with arithmetic in such finite sets. Example 1.5. We consider the set of the nine numbers: {0, 1, 2, 3, 4, 5, 6, 7, 8} We can do regular arithmetic as long as the results are smaller than 9. For instance: 2 · 3 = 6 4 + 4 = 8
16 1 Introduction to Cryptography and Data Security But what about 8 + 4? Now we try the following rule: Perform regular integer arith- metic and divide the result by 9. We then consider only the remainder rather than the original result. Since 8 + 4 = 12, and 12/9 has a remainder of 3, we write: 8 + 4 ≡ 3 mod 9 We now introduce an exact definition of the modulo operation. Definition 1.4.1 Modulo Operation Let a, r, m ∈ Z (where Z is a set of all integers) and m > 0. We write a ≡ r mod m if m divides a − r. m is called the modulus and r is called the remainder. There are implications from this definition that go beyond the casual rule “divide by the modulus and consider the remainder.” We discuss these in the following. Computation of the Remainder It is always possible to write a ∈ Z, such that a = q · m + r for 0 ≤ r < m (1.1) Since a − r = q · m, i.e., m divides a − r, we can now write: a ≡ r mod m. Note that r ∈ {0, 1, 2, . . . , m − 1}. Example 1.6. Let a = 42 and m = 9. Then 42 = 4 · 9 + 6 and therefore 42 ≡ 6 mod 9. The Remainder Is Not Unique It is somewhat surprising that for every given modulus m and number a, there are (infinitely) many valid remainders. Let’s look at another example: Example 1.7. We want to reduce 12 modulo 9. Here are several results that are cor- rect according to the definition:
1.4 Modular Arithmetic and More Historical Ciphers 17 12 ≡ 3 mod 9, 3 is a valid remainder since 9|(12 − 3) 12 ≡ 21 mod 9, 21 is a valid remainder since 9|(12 − 21) 12 ≡ −6 mod 9, −6 is a valid remainder since 9|(12 − (−6)) where the “x|y” means “x divides y”. There is a system behind this behavior. The set of numbers: {. . . , −24, −15, −6, 3, 12, 21, 30, . . .} form what is called an equivalence class. There is a total of nine equivalence classes for the modulus 9: {. . . , −27, −18, −9, 0, 9, 18, 27, . . .} {. . . , −26, −17, −8, 1, 10, 19, 28, . . .} ... {. . . , −19, −10, −1, 8, 17, 26, 35, . . .} We note that every integer, i.e., every number without decimal places from minus infinity to plus infinity, is a member in one of these equivalence classes. All Members of a Given Equivalence Class Behave Equivalently For a given modulus m, it does not matter which element from a class we choose for a given computation. This property of equivalence classes has major practical implications. If we have involved computations with a fixed modulus — which is usually the case in cryptography — we are free to choose the class element that results in the easiest computation. Let’s look first at an example. Example 1.8. The core operation in many practical public-key schemes is an expo- nentiation of the form xe mod m, where x, e, m are very large integers, say, 2048 bits each. Using a toy-size example, we can demonstrate two ways of doing modular ex- ponentiation. We want to compute 38 mod 7. The first method is the straightforward approach, and for the second one we switch within the equivalence class. 1. Na¨ıve method: We compute 38 = 6561 ≡ 2 mod 7, since 6561 = 937 · 7 + 2. Note that we obtain the fairly large intermediate result 6561 even though we know that our final result cannot be larger than 6. 2. Here is a much smarter method: First we perform two partial exponentiations: 38 = 34 · 34 = 81 · 81 We can now replace the intermediate results 81 by another member of the same equivalence class. The smallest positive member modulo 7 in the class is 4 (since 81 = 11 · 7 + 4). Hence:
18 1 Introduction to Cryptography and Data Security 38 = 81 · 81 ≡ 4 · 4 = 16 mod 7 From here we obtain the final result easily as 16 ≡ 2 mod 7. Note that we could perform the second method without a pocket calculator since the numbers never become larger than 81. For the first method, on the other hand, dividing 6561 by 7 is mentally already a bit challenging. As a general rule we should remember that it is almost always of computational advantage to apply the modulo reduction as soon as we can in order to keep the numbers small. Of course, the final result of any modulo computation is always the same, no matter how often we switch back and forth within equivalence classes. Which Remainder Do We Choose? By agreement, we usually choose r in Equation (1.1) such that: 0 ≤ r ≤ m − 1 However, mathematically it does not matter which member of an equivalent class we use. 1.4.2 Integer Rings After studying the properties of modulo reduction we are now ready to define in more general terms a structure that is based on modulo arithmetic. Let’s look at the mathematical construction that we obtain if we consider the set of integers from zero to m − 1 together with the operations addition and multiplication. Definition 1.4.2 Ring The integer ring Zm consists of: 1. The set Zm = {0, 1, 2, . . . , m − 1} 2. Two operations “+” and “·” for all a, b ∈ Zm such that: 1. a + b ≡ c mod m (c ∈ Zm) 2. a · b ≡ d mod m (d ∈ Zm)
1.4 Modular Arithmetic and More Historical Ciphers 19 Let’s first look at an example of a small integer ring. Example 1.9. Let m = 9, i.e., we are dealing with the ring Z9 = {0, 1, 2, 3, 4, 5, 6, 7, 8}. Here are two simple computations in this ring: 6 + 8 = 14 ≡ 5 mod 9 6 · 8 = 48 ≡ 3 mod 9 More about rings and finite fields, which are related to rings, is discussed in Section 4.3. At this point, the following properties of rings are important: We can add and multiply any two numbers from the set and the result is always in the ring. A ring is said to be closed. Addition and multiplication are associative, i.e., a + (b + c) = (a + b) + c and a · (b · c) = (a · b) · c, for all a, b, c ∈ Zm. Addition is commutative, i.e., a + b = b + a, for all a, b ∈ Zm. There is the neutral element 0 with respect to addition, i.e., for every element a ∈ Zm it holds that a + 0 ≡ a mod m. For any element a in the ring, there is always the negative element −a such that a + (−a) ≡ 0 mod m, i.e., the additive inverse always exists. There is the neutral element 1 with respect to multiplication, i.e., for every ele- ment a ∈ Zm it holds that a · 1 ≡ a mod m. The multiplicative inverse exists only for some, but not for all, elements. Let a ∈ Z. The inverse a−1 is defined such that a · a−1 ≡ 1 mod m If an inverse exists for a, we can divide by this element since b/a ≡ b · a−1 mod m. Another ring property is that a · (b + c) = (a · b) + (a · c) for all a, b, c ∈ Zm, i.e., the distributive law holds. In summary, roughly speaking, we can say that the ring Zm is the set of integers {0, 1, 2, . . . , m − 1} in which we can add, subtract, multiply and sometimes divide. One issue that is worth discussing is the multiplicative inverse. It takes some ef- fort to find the inverse (usually employing the extended Euclidean algorithm, which is introduced in Section 6.3). However, there is an easy way of telling whether an inverse for a given element a exists or not: An element a ∈ Z has a multiplicative inverse a−1 if and only if gcd(a, m) = 1, where gcd is the greatest common divisor, i.e., the largest integer that divides both numbers a and m. The fact that two numbers have a gcd of 1 is of importance in number theory, and there is a special name for it: If gcd(a, m) = 1, then a and m are said to be relatively prime or coprime.
20 1 Introduction to Cryptography and Data Security Example 1.10. Let’s see whether the multiplicative inverse of 15 exists in Z26. Be- cause gcd(15, 26) = 1 the inverse must exist. (In fact, the inverse is 7 since 7·15 ≡ 1 mod 26.) On the other hand, since gcd(14, 26) = 2 6 = 1 the multiplicative inverse of 14 does not exist in Z26. As mentioned earlier, the ring Zm, and thus integer arithmetic with the modulo operation, is of central importance in modern public-key cryptography. In practice, the integers involved have a length of 256–4096 bits so that we need ways to perform modular arithmetic with such large numbers efficiently. 1.4.3 Shift Cipher (or Caesar Cipher) We now introduce another historical cipher, the shift cipher. It is actually a special case of the substitution cipher and has a very elegant mathematical description. The shift cipher itself is extremely simple: We simply shift every plaintext letter by a fixed number of positions in the alphabet. For instance, if we shift by 3 posi- tions, A would be substituted by d, B by e, etc. The only problem arises towards the end of the alphabet: What should we do with X, Y, Z? As you might have guessed, they should “wrap around”. That means X should become a, Y should be- come b, and Z is replaced by c. (In light of this rule, a more accurate name for the shift cipher would be “rotation cipher” but this name is rarely used.) Allegedly, Julius Caesar used this cipher with a three-position shift. The shift cipher also has an elegant description using modular arithmetic. For the mathematical representation of the cipher, the letters of the alphabet are encoded as numbers, as depicted in Table 1.3. Table 1.3 Encoding of letters for the shift cipher A B C D E F G H I J K L M 0 1 2 3 4 5 6 7 8 9 10 11 12 N O P Q R S T U V W X Y Z 13 14 15 16 17 18 19 20 21 22 23 24 25 Both the plaintext letters and the ciphertext letters are now elements of the ring Z26. Also, the key, i.e., the number of shift positions, is in Z26 since more than 26 shifts would not make sense (27 shifts would be the same as 1 shift, etc.). The encryption and decryption of the shift cipher are as follows.
1.4 Modular Arithmetic and More Historical Ciphers 21 Definition 1.4.3 Shift Cipher Let x, y, k ∈ Z26. Encryption: ek(x) ≡ x + k mod 26 Decryption: dk(y) ≡ y − k mod 26 Example 1.11. Let the key be k = 17, and the plaintext is: ATTACK = x1, x2, . . . , x6 = 0, 19, 19, 0, 2, 10 The ciphertext is then computed as y1, y2, . . . , y6 = 17, 10, 10, 17, 19, 1 = rkkrtb As you can guess from the discussion of the substitution cipher earlier in this book, the shift cipher is not secure at all. There are two ways of attacking it: 1. Since there are only 26 different keys (shift positions), one can easily launch a brute-force attack by trying to decrypt a given ciphertext with all possible 26 keys. If the resulting plaintext is readable text, you have found the key. 2. As for the substitution cipher, one can also use letter frequency analysis. The attack works even better for the shift cipher than for the substitution cipher. As soon as the attacker has discovered the ciphertext letter for one plaintext letter, he/she knows the number of shifts and thus has the key. 1.4.4 Affine Cipher We try now to improve the shift cipher by generalizing the encryption function. Recall that the actual encryption of the shift cipher was the addition of the key yi ≡ xi + k mod 26. The affine cipher encrypts by multiplying the plaintext by one part of the key followed by addition of another part of the key. Definition 1.4.4 Affine Cipher Let x, y, a, b ∈ Z26. Encryption: ek(x) = y ≡ a · x + b mod 26 Decryption: dk(y) = x ≡ a−1 · (y − b) mod 26 with the key: k = (a, b), which has the restriction: gcd(a, 26) = 1.
22 1 Introduction to Cryptography and Data Security The decryption is easily derived from the encryption function: a · x + b ≡ y mod 26 a · x ≡ (y − b) mod 26 x ≡ a−1 · (y − b) mod 26 The restriction gcd(a, 26) = 1 stems from the fact that the key parameter a needs to be inverted for decryption. We recall from Section 1.4.2 that an element a and the modulus must be relatively prime for the inverse of a to exist. Thus, a must be in the set: a ∈ {1, 3, 5, 7, 9, 11, 15, 17, 19, 21, 23, 25} (1.2) But how do we find a−1? For now, we can simply compute it by trial and error: For a given a we simply try all possible values a−1 until we obtain: a · a−1 ≡ 1 mod 26 For instance, if a = 3, then a−1 = 9 since 3 · 9 = 27 ≡ 1 mod 26. Note that a−1 also always fulfills the condition gcd(a−1, 26) = 1 since the inverse of a−1 always exists. In fact, the inverse of a−1 is a itself. Hence, for the trial-and-error determination of a−1 one only has to check the values given in Equation (1.2). Example 1.12. Let the key be k = (a, b) = (9, 13), and the plaintext be ATTACK = x1, x2, . . . , x6 = 0, 19, 19, 0, 2, 10. The ciphertext is computed as y1, y2, . . . , y6 = 13, 2, 2, 13, 5, 25 = nccnfz For decryption, the inverse of a needs to be determined, which turns out to be a−1 = 3. Is the affine cipher secure? No! The key space is only a bit larger than in the case of the shift cipher: key space = (#values for a) · (#values for b) = 12 · 26 = 312 A key space with 312 elements can, of course, still be searched exhaustively, i.e., brute-force attacked, in a fraction of a second with any PC. In addition, the affine cipher has the same weakness as the shift and substitution cipher: The mapping between plaintext letters and ciphertext letters is fixed. Hence, it can also be broken with letter frequency analysis. The remainder of this book deals with strong cryptographic algorithms which are of practical relevance.
1.5 Discussion and Further Reading 23 1.5 Discussion and Further Reading This book addresses practical aspects of cryptography and data security and is in- tended to be used as an introduction; it is suited for classroom use, distance learning and self-study. At the end of each chapter, we provide a discussion section in which we briefly describe topics for readers interested in further study of the material. Cryptography vs. Cybersecurity vs. Safety and Reliability As mentioned at the very beginning of the book, cryptography is part of the broader fields of cyber- security and IT security, where it is difficult to have a clear distinction between those two latter terms. In fact, there exist many definitions for IT- and cybersecu- rity. Traditionally, those terms were often described as dealing with “assurance of the confidentiality, integrity and availability of information”, sometimes referred to as the CIA triad. However, in addition to these three basic security goals, there are often additional ones, including authenticity, accountability, non-repudiation and re- liability. More about security services can be found in Section 10.1.3 of this book. It is important to bear in mind that cryptography, IT- and cybersecurity all deal with the protecting of information systems against malicious human actors, to which we refer as attackers or adversaries in this book. In contrast, technical safety4 is con- cerned with protection against dangers such as random failures that arise during the regular use of technical systems. For instance, when driving a car, we want to ensure that the brakes and the steering don’t fail — otherwise it would be unsafe. In order to achieve such technical safety, systems must be reliable. In contrast to security, safety and reliability are primarily not concerned with failure due to malicious ac- tors but due to (random) technical failures. Even though reliability and security are partially interdependent, they involve different aspects of protecting systems. In order to approach the problem of IT security systematically, several general frameworks exist. They typically follow a holistic approach by taking all security- relevant factors into account. Such an approach requires that assets and correspond- ing security needs have to be defined, and that the attack potential and possible attack paths must be evaluated. Finally, adequate countermeasures have to be spec- ified in order to realize an appropriate level of security for a particular application or environment. There are numerous standards that can be used for evaluation and help to define a secure system. Among the more prominent ones are ISO/IEC 27001 for Information Security Management Systems (ISMS), the Common Criteria for Information Technology Security Evaluation [75] and FIPS PUBS [116]. In some industries, standards help to establish a more domain-specific approach towards IT security, e.g., ISO/IEC 62443 for industrial communication networks or ISO/SAE 21434 for cybersecurity engineering for road vehicles [147]. Moreover, frameworks such as the NIST framework for improving the IT security in critical infrastructures exist [29]. Historical Ciphers and Kerckhoffs’ Principle This chapter introduced a few his- torical ciphers. However, there are many, many more, ranging from ciphers in an- 4 We note that safety is also used in non-technical contexts, e.g., food safety.
24 1 Introduction to Cryptography and Data Security cient times to WWII encryption methods. To readers who wish to learn more about historical ciphers and the role they played over the centuries, the books by Bauer [30], Kahn [156], Singh [237] and Wrixon [254] are recommended. Besides mak- ing fascinating bedtime reading, these books help one to understand the role that military and diplomatic intelligence played in shaping world history. They also help to show modern cryptography in a larger context. Auguste Kerckhoffs was a Dutch cryptographer and linguist in the second half of the nineteenth century. He observed that cryptography is often used incorrectly in practice and postulated six principles in 1883, given below. What’s today widely known as Kerckhoffs’ Principle is actually the second one from the list. The system should be, if not theoretically unbreakable, unbreakable in practice. The design of a system should not require secrecy, and compromise of the system should not inconvenience the correspondents. The key should be memorable without notes and should be easily changeable. The cryptograms should be transmittable by telegraph. The apparatus or documents should be portable and operable by a single person. The system should be easy, neither requiring knowledge of a long list of rules nor involving mental strain. It is notable that several of the principles deal with the use of cryptography, as opposed to the technical aspects of ciphers, a fact that was only relatively recently observed by Sasse [226]. It was only in the 1990s that usable security emerged as a proper research discipline within the scientific community. Modular Arithmetic The mathematics introduced in this chapter, modular arith- metic, belongs to the field of number theory. This is a fascinating subject area which was, unfortunately, historically viewed as a “branch of mathematics with- out applications”. Thus, it is rarely taught outside mathematics curricula. There is a wealth of books on number theory. Among the classic introductory books are refer- ences [204, 222]. A particularly accessible book written for non-mathematications is reference [236]. Provable Security Due to our focus on practical cryptography, this book omits most aspects related to the theoretical foundations of cryptographic algorithms and protocols. One of the foundations of theoretical cryptography builds on the belief that any cryptographic scheme should be accompanied by a rigorous mathematical proof of its security (“security proof”) under a well-defined and reasonable crypto- graphic hardness assumption. Examples of such hardness assumptions include the assumption that computing discrete logarithms over certain prime-order groups is difficult, the assumption that finding a small vector in a high-dimensional lattice is difficult, or the assumption that finding a collision in a concrete hash function is difficult. This concept is called provable security.5 Informally, “provable security” 5 The term “provable security” may be slightly misleading since it does not provide unconditional proofs in a mathematical sense. It rather reduces a protocol’s security to a well-defined mathe- matical hardness assumption. Some cryptographers therefore prefer to use the term “reductionist security” instead.
1.5 Discussion and Further Reading 25 is achieved for a given cryptographic scheme when one provides (i) an algorith- mic description of the cryptographic scheme; (ii) a formulation of a rigorous and precise definition of the adversary’s capacities and goal (security model); and (iii) a mathematical proof that the proposed scheme meets its security goal, assuming some standard cryptographic assumption holds true. Point (iii) is remarkable and deserves more attention. This mathematical proof shows formally that the only way to break the scheme (within the defined security model) is to attack the underly- ing cryptographic assumption. The proof holds for all possible attacks under the same assumptions, even the ones we could not envision at the time of designing the scheme. The standard references for provable security are the textbooks by Katz/Lindell [158] and Goldreich [123, 124]. Also recommended is the more recent online book by Rosulek [223]. A few times this book also touches upon provable security, for instance the re- lationship between Diffie–Hellman key exchange and the Diffie–Hellman problem (cf. Section 8.4), the block cipher-based hash functions in Section 11.3.1, the secu- rity of the HMAC message authentication scheme in Section 13.2, or the security of lattice-based cryptography based on the conjectured intractability of the shortest- vector problem in Section 12.2. Advanced Cryptographic Schemes There are many advanced cryptographic con- structions that go beyond the symmetric and asymmetric ciphers that are the main topic of this book. In the following we sketch some of the more important examples of advanced cryptography. Homomorphic encryption allows computation on encrypted data, i.e., on cipher- text, without first decrypting. A major application scenario is cloud computing, where a user has a massive amount of data in encrypted form in the cloud. If the data is, for instance, a large customer database, the user might be interested in download- ing some customer records that fulfill certain criteria. The challenge is to perform such a search on the ciphertext. It is relatively easy to construct partially homo- morphic encryption schemes, which are constructions that allow one mathematical operation to be performed on the ciphertext, typically multiplication or addition. In fact, the two popular asymmetric encryption schemes RSA (cf. Chapter 7) and Elga- mal (cf. Section 8.5) are partially homomorphic. Unfortunately, one mathematical operation is not sufficient for the majority of practical applications. For a long time finding a fully homomorphic encryption scheme that allows arbitrary operations was considered the holy grail of cryptography. The first such scheme was proposed by Gentry [122] in 2009, which is based on lattices (cf. Section 12.2). This original sys- tem was quite impractical but since then numerous improvements have taken place. At the time of writing, many competing schemes exist and use in practice is within reach. This topic also relevant for the training of machine learning algorithms. Another advanced cryptographic scheme is multiparty computation (MPC), also known as secure multiparty computation. With MPC, several parties provide input values and jointly compute a function from the inputs. The interesting part is that when the protocol is completed the participants know only their own input and the
26 1 Introduction to Cryptography and Data Security answer but nothing about the inputs of the other participants. A standard example is a situation where three people want to find out what the highest salary in the group is without revealing the individual salaries. Another application is determining the outcome of an election, that is electronic voting, or the highest bid in an auction based on encrypted data. The general theory of MPC was proposed in the late 1980s but it took more than 20 years before the first practical application started to emerge. A good reference source is [80]. Related to multiparty computation is secret sharing. The idea of (general) secret sharing is that out of n participants t must collaborate to compute a secret, e.g., a cryptographic key. A real-world scenario is that at least 2 out of 3 managers of a bank must get together to generate the secret code for opening a safe. Secret sharing was proposed independently by Shamir and Blakley in 1979 [229, 54]. Zero-knowledge proofs are concerned with proving certain knowledge to another party without revealing the secret. They were originally motivated for authentication without revealing a password or key. There are many other applications such as anonymous payment schemes. Zero-knowledge proofs were originally proposed by Goldwasser, Micali and Rackoff [125]. Other advanced cryptographic constructions include identity-based encryption, attribute-based encryption, functional encryption and proxy-reencryption. Research Community and General References Even though cryptography has matured considerably since the 1970s, it is still a relatively young field compared to other scientific disciplines, and every year brings many new developments and dis- coveries. Many research results are published at the eight main events organized by the International Association for Cryptologic Research (IACR). The proceedings of the three IACR conferences Crypto, Eurocrypt and Asicacrypt, the four more spe- cific area conferences Cryptographic Hardware and Embedded Systems6 (CHES), Fast Software Encryption (FSE), Public Key Cryptography (PKC) and Theoretical Cryptograpy Conference (TCC), as well as the Real World Cryptography (RWC) symposium are excellent sources for tracking recent developments in the field of cryptology. There are four top conferences in the broader field of computer secu- rity (of which cryptography is one aspect): the IEEE Symposium on Security and Privacy (IEEE S&P), the ACM Conference on Computer and Communications Se- curity (CCS), the USENIX Security Symposium and the Network and Distributed System Security Symposium (NDSS). It should be stressed that in cryptography as well as in computer security there are many, many more conferences and workshops, many of which are also of very high quality. There are several good books on cryptography. A classic, if somewhat dated, book is Applied Cryptography [227] by Schneier published in 1994, which helped to popularize modern cryptography. A more recent book, which makes an excellent addition to the book at hand, is Serious Cryptography: A Practical Introduction to Modern Encryption by Aumasson [20]. With respect to reference sources, the Handbook of Applied Cryptography by Menezes, van Oorschot and Vanstone [189] and the Encyclopedia of Cryptography and Security [246] can be recommended. An 6 CHES was co-founded by one of the authors of this book.
1.6 Lessons Learned 27 excellent reference for the much broader field of security engineering is Anderson’s Security Engineering: A Guide to Building Dependable Distributed Systems [12]. 1.6 Lessons Learned Never ever develop your own cryptographic algorithm unless you have a team of experienced cryptanalysts checking your design. Do not use unproven cryptographic algorithms (i.e., symmetric ciphers, asym- metric ciphers, hash functions) or unproven protocols. Attackers always look for the weakest point of a cryptosystem. For instance, a large key space by itself is no guarantee of a cipher being secure; the cipher might still be vulnerable against analytical attacks. Key lengths for symmetric algorithms in order to thwart exhaustive key-search attacks are: 64 bits: insecure except for data with extremely short-term value. 112–128 bits: long-term security of several decades, including attacks by in- telligence agencies unless they possess quantum computers. Based on our cur- rent knowledge, attacks are only feasible with quantum computers (which do not exist but might become reality in 1–2 decades). 256 bits: as above, but possibly secure against attacks by quantum computers. Modular arithmetic is a tool for expressing historical encryption schemes, such as the affine cipher, in a mathematically elegant way and provides the fundamental basis for many modern cryptographic schemes.
28 1 Introduction to Cryptography and Data Security Problems 1.1. The ciphertext below was encrypted using a substitution cipher. Decrypt the ci- phertext without knowledge of the key. lrvmnir bpr sumvbwvr jx bpr lmiwv yjeryrkbi jx qmbm wi bpr xjvni mkd ymibrut jx irhx wi bpr riirkvr jx ymbinlmtmipw utn qmumbr dj w ipmhh but bj rhnvwdmbr bpr yjeryrkbi jx bpr qmbm mvvjudwko bj yt wkbrusurbmbwjk lmird jk xjubt trmui jx ibndt wb wi kjb mk rmit bmiq bj rashmwk rmvp yjeryrkb mkd wbi iwokwxwvmkvr mkd ijyr ynib urymwk nkrashmwkrd bj ower m vjyshrbr rashmkmbwjk jkr cjnhd pmer bj lr fnmhwxwrd mkd wkiswurd bj invp mk rabrkb bpmb pr vjnhd urmvp bpr ibmbr jx rkhwopbrkrd ywkd vmsmlhr jx urvjokwgwko ijnkdhrii ijnkd mkd ipmsrhrii ipmsr w dj kjb drry ytirhx bpr xwkmh mnbpjuwbt lnb yt rasruwrkvr cwbp qmbm pmi hrxb kj djnlb bpmb bpr xjhhjcwko wi bpr sujsru msshwvmbwjk mkd wkbrusurbmbwjk w jxxru yt bprjuwri wk bpr pjsr bpmb bpr riirkvr jx jqwkmcmk qmumbr cwhh urymwk wkbmvb 1. Compute the relative frequency of all letters A...Z in the ciphertext. You may want to use a tool such as the open-source program CrypTool [82] for this task. However, a paper and pencil approach is also doable. 2. Decrypt the ciphertext with the help of the relative letter frequency of the English language (see Table 1.1 in Section 1.2.2). Note that the text is relatively short and that the letter frequencies in it might not perfectly align with that of general English language from the table. 3. Who wrote the text? 1.2. We received the following ciphertext which was encoded with a shift cipher: xultpaajcxitltlxaarpjhtiwtgxktghidhipxciwtvgtpilpit ghlxiwiwtxgqadds. 1. Perform an attack against the cipher based on a letter frequency count: How many letters do you have to identify through a frequency count to recover the key? What is the cleartext? 2. Who wrote this message? 1.3. We consider the long-term security of the Advanced Encryption Standard (AES) with a key length of 128 bits with respect to exhaustive key-search attacks. AES is perhaps the most widely used symmetric cipher at this time. 1. Assume that an attacker has special-purpose hardware chips (also known as ASICs, or application-specific integrated circuits) that check 5 · 108 keys per
1.6 Problems 29 second, and she has a budget of $1 million. One ASIC costs $50, and we as- sume 100% overhead for integrating the ASIC (manufacturing the printed circuit boards, power supply, cooling, etc.). How many ASICs can we run in parallel with the given budget? How long does an average key search take? Relate this time to the age of the Universe, which is about 1010 years. 2. We try now to take advances in computer technology into account. Predicting the future tends to be tricky but the estimate usually applied is Moore’s law, which states that the computing power doubles every 18 months while the costs of integrated circuits stay constant. How many years do we have to wait until a key-search machine can be built to break AES with 128 bits with an average search time of 24 hours? Again, assume a budget of $1 million (do not take inflation into account). 1.4. We now consider the relation between passwords and key size. For this purpose we consider a cryptosystem where the user enters a key in the form of a password. 1. Assume a password consisting of 8 letters, where each letter is encoded with the ASCII code (7 bits per character, i.e., 128 possible characters). What is the size of the key space which can be constructed by such passwords? 2. What is the corresponding key length in bits? 3. Assume that most users use only the 26 lowercase letters from the alphabet in- stead of the full 7 bits of the ASCII-encoding. What is the corresponding key length in bits in this case? 4. At least how many characters are required for a password in order to generate a key length of 128 bits in case of letters consisting of a. 7-bit characters? b. 26 lowercase letters from the alphabet? 1.5. In case of a brute-force attack, we have to search the entire key space of a cipher. To prevent such a search from being successful, the key space must be sufficiently large. It is crucial to observe that the key space grows exponentially with the key length in bits. With this problem we want to get a better understanding of such an exponential growth. According to an anecdote, the inventor of chess asked the king for a humble reward in the form of grains of rice: On the first field of the chess board, the king should put one grain of rice, on the second field two grains of rice, on the third field four grains etc. 1. How many grains of rice are on the last field of the chess board? 2. A single grain of rice has a weight of approximately 0.03 g. What is the total weight of all grains on the board? Compare the total weight with the worldwide yield of approximately 480 million tons per year. Now, let us consider a piece of paper that is repeatedly folded. The thickness of the paper increases exponentially: It has twice the thickness if folded once, four times the thickness if folded twice etc. For the following tasks, we assume a piece of paper which is 0.1 mm thick.
30 1 Introduction to Cryptography and Data Security 3. How thick is the paper after 10 folding steps? 4. How often do we need to fold it to obtain a thickness of 1 km? 5. How often do we need to fold it to obtain the distance from the Earth to the Moon (384,400 km)? 6. How often do we need to fold it to obtain the distance of one light year, i.e., 9.46 · 1015 km? Remark: Obviously, folding a piece of paper that often will not work out very well in practice. 1.6. In this problem we consider the difference between end-to-end encryption (E2EE) and more classical approaches to encrypting when communicating over a channel that consists of multiple parts. E2EE is widely used, e.g., in instant messag- ing services such as WhatsApp or Signal. The idea behind this is that encryption and decryption are performed by the two users who communicate and all parties eaves- dropping on the communication link cannot read (or meaningfully manipulate) the message. In the following we assume that each individual encryption with the cipher e() is secure, i.e., the cryptographic algorithm cannot be broken by an adversary. First we look at the communication between two smartphones without end-to-end encryp- tion, shown in Figure 1.7. Encryption and/or decryption happen three times in this setting: Between Alice and base station A (air link), between base stations A and B (through the internet), and between base station B and Bob (again, air link). Fig. 1.7 Communication without E2EE 1. Describe which of the following attackers can read (and meaningfully manipu- late) messages. a. A hacker who can listen to (and alter) messages on the air link between Alice and her base station. b. The mobile operator who runs and controls base station A. c. A national law enforcement agency that has power over the mobile operator and gains access to base station A or B. d. An intelligence agency of a foreign country that can wiretap any internet com- munication. e. The mobile operator who runs and controls base station B.
1.6 Problems 31 f. A hacker who can listen to (and alter) messages on the air link between Bob and his base station. We now look at the same communication system but this time Alice and Bob use E2EE, cf. Figure 1.8 Fig. 1.8 Communication with E2EE 2. Describe which of the following attackers can read (and meaningfully manipu- late) messages in the communication systems with E2EE. a. A hacker who can listen to (and alter) messages on the air link between Alice and her base station. b. The mobile operator who runs and controls base station A. c. A national law enforcement agency that has power over the mobile operator and gains access to base station A or B. d. An intelligence agency of a foreign country that can wiretap any internet com- munication. e. The mobile operator who runs and controls base station B. f. A hacker who can listen to (and alter) messages on the air link between Bob and his base station. 1.7. As we learned in this chapter, modular arithmetic is the basis of many cryp- tosystems. We will now provide a number of exercises that help us get familiar with modular computations. Let’s start with an easy one: Compute the following result without a calculator. 1. 15 · 29 mod 13 2. 2 · 29 mod 13 3. 2 · 3 mod 13 4. −11 · 3 mod 13 The results should be given in the range from 0, 1, . . . , modulus-1. Briefly describe the relation between the different parts of the problem.
32 1 Introduction to Cryptography and Data Security 1.8. Compute without a calculator: 1. 1/5 mod 13 2. 1/5 mod 7 3. 3 · 2/5 mod 7 1.9. We consider the ring Z4. Construct a table that describes the addition of all elements in the ring with each other in the following form: + 0 1 2 3 0 0 1 2 3 1 1 2 · · · 2 · · · 3 1. Construct the multiplication table for Z4. 2. Construct the addition and multiplication tables for Z5. 3. Construct the addition and multiplication tables for Z6. 4. There are elements in Z4 and Z6 without a multiplicative inverse. Which ele- ments are these? Why does a multiplicative inverse exist for all nonzero elements in Z5? 1.10. What is the multiplicative inverse of 5 in Z11, Z12, and Z13? You can do a trial-and-error search using a calculator or a PC. With this simple problem we want now to stress the fact that the inverse of an integer in a given ring depends completely on the ring considered. That is, if the modulus changes, the inverse changes. Hence, it doesn’t make sense to talk about an inverse of an element unless it is clear what the modulus is. This fact is crucial for the RSA cryptosystem, which is introduced in Chapter 7. The extended Euclidean algorithm, which can be used for computing inverses efficiently, is introduced in Section 6.3. 1.11. Compute x as far as possible without a calculator. Where appropriate, make use of a smart decomposition of the exponent as shown in the example in Sec- tion 1.4.1: 1. x ≡ 32 mod 13 2. x ≡ 72 mod 13 3. x ≡ 310 mod 13 4. x ≡ 7100 mod 13 5. 7x ≡ 11 mod 13 The last problem is called a discrete logarithm and points to a hard problem which we discuss in Chapter 8. The security of many public-key schemes is based on the hardness of solving the discrete logarithm for large numbers, e.g., with more than 2000 bits.
1.6 Problems 33 1.12. Find all integers n with 0 ≤ n < m that are relatively prime to m for m = 4, 5, 9, 26. We denote the number of integers n which fulfill the condition by φ (m), e.g., φ (3) = 2. This function is called “Euler’s phi function”. What is φ (m) for m = 4, 5, 9, 26? More on Euler’s phi function will be said in Section 6.3. 1.13. This problem deals with the affine cipher where the key is given as a = 7 and b = 22. 1. Decrypt the text below: falszztysyjzyjkywjrztyjztyynaryjkyswarztyegyyj 2. Who wrote the line? 1.14. We want to extend the affine cipher from Section 1.4.4 such that we can en- crypt and decrypt messages written with the full German alphabet. The German alphabet consists of the English one together with the three umlauts, ¨A, ¨O, ¨U, and the (even stranger) “sharp S” character ß. We use the following mapping from letters to integers: A ↔ 0 B ↔ 1 C ↔ 2 D ↔ 3 E ↔ 4 F ↔ 5 G ↔ 6 H ↔ 7 I ↔ 8 J ↔ 9 K ↔ 10 L ↔ 11 M ↔ 12 N ↔ 13 O ↔ 14 P ↔ 15 Q ↔ 16 R ↔ 17 S ↔ 18 T ↔ 19 U ↔ 20 V ↔ 21 W ↔ 22 X ↔ 23 Y ↔ 24 Z ↔ 25 ¨A ↔ 26 ¨O ↔ 27 ¨U ↔ 28 ß ↔ 29 1. What are the encryption and decryption equations for the cipher? 2. How large is the key space of the affine cipher for this alphabet? 3. The following ciphertext was encrypted using the key (a = 17, b = 1). What is the corresponding plaintext? ¨a u ß w ß 4. From which village does the plaintext come? 1.15. We consider an attack scenario where the adversary Oscar manages to provide Alice with a few pieces of plaintext that she encrypts. Show how Oscar can break the affine cipher by using two pairs of plaintext–ciphertext, (x1, y1) and (x2, y2). What is the condition for choosing x1 and x2? Remark: In practice, this chosen-plaintext attack is often possible depending on the application, e.g., if Alice is a web server that encrypts and returns messages that are sent to her. 1.16. An obvious approach to increase the security of a symmetric algorithm is to apply the same cipher twice, i.e., y = ek2(ek1(x)) As is often the case in cryptography, things can be tricky and results are often dif- ferent from the expected or desired ones. In this problem we show that a double
34 1 Introduction to Cryptography and Data Security encryption with the affine cipher is only as secure as single encryption! Assume two affine ciphers ek1 ≡ a1x + b1 mod 26 and ek2 ≡ a2x + b2 mod 26. 1. Show that there is a single affine cipher ek3 ≡ a3x + b3 mod 26 which performs exactly the same encryption (and decryption) as the combination ek2(ek1(x)). 2. Find the values for a3, b3 when a1 = 3, b1 = 5 and a2 = 11, b2 = 7. 3. To verify your solution, (1) encrypt the letter K with ek1 and the result with ek2, and (2) encrypt the letter K with ek3. 4. Briefly describe what happens if an exhaustive key-search attack is applied to a double-encrypted affine ciphertext. Is the effective key space increased? Remark: The issue of multiple encryption is of great practical importance in the case of the Data Encryption Standard (DES), for which multiple encryption (in par- ticular, triple encryption) does increase security considerably, cf. Section 5.3.2. 1.17. We already know that the substitution cipher and the shift cipher can easily be broken in little time. Let us now consider an extension of the shift cipher, namely the Vigen`ere cipher (named after Blaise de Vigen`ere). Instead of using a single key k for the shift, it uses l different shifts that are derived from a secret code word c. The code word consists of l letters and has the form c = (c0, c1, . . . , cl−1). Each letter ci corresponds to a number 0, . . . , 25, which is given by its position in the alphabet. These numbers are the l shift positions, which we denote by (k0, k1, . . . kl−1). Encryption (and decryption) work as follows: The first plaintext letter x0 is cycli- cally shifted by k0 positions, the second plaintext x1 by k1 positions and so on, until plaintext letter xl−1 is shifted by kl−1 positions. From now on, the shift sequence repeats, i.e., plaintext xl is again shifted by k0 positions, the next plaintext by k1 positions and so on. This process is expressed as: y j ≡ x j + k( j mod l) mod 26 where x j denotes the j-th letter of the plaintext x = (x0, x1, . . .). Since the cipher uses many ciphertext alphabets, it is called a polyalphabetic cipher. 1. Assume the code word is given as c = JAMAIKA of size l = 7. Transform the code word into the corresponding encryption keys ki. You can use Table 1.4 for this task. 2. Use the table to encrypt the word x = CODEBREAKERS with the Vigen`ere ci- pher. For each plaintext letter, choose the row with the corresponding shift value in the leftmost column # and look up the shifted version of the plaintext. 3. What do you think about the security of the Vigen`ere cipher? Propose an attack.
1.6 Problems 35 # A B C D E F G H I J K L M N O P Q R S T U V W X Y Z 0 A B C D E F G H I J K L M N O P Q R S T U V W X Y Z 1 B C D E F G H I J K L M N O P Q R S T U V W X Y Z A 2 C D E F G H I J K L M N O P Q R S T U V W X Y Z A B 3 D E F G H I J K L M N O P Q R S T U V W X Y Z A B C 4 E F G H I J K L M N O P Q R S T U V W X Y Z A B C D 5 F G H I J K L M N O P Q R S T U V W X Y Z A B C D E 6 G H I J K L M N O P Q R S T U V W X Y Z A B C D E F 7 H I J K L M N O P Q R S T U V W X Y Z A B C D E F G 8 I J K L M N O P Q R S T U V W X Y Z A B C D E F G H 9 J K L M N O P Q R S T U V W X Y Z A B C D E F G H I 10 K L M N O P Q R S T U V W X Y Z A B C D E F G H I J 11 L M N O P Q R S T U V W X Y Z A B C D E F G H I J K 12 M N O P Q R S T U V W X Y Z A B C D E F G H I J K L 13 N O P Q R S T U V W X Y Z A B C D E F G H I J K L M 14 O P Q R S T U V W X Y Z A B C D E F G H I J K L M N 15 P Q R S T U V W X Y Z A B C D E F G H I J K L M N O 16 Q R S T U V W X Y Z A B C D E F G H I J K L M N O P 17 R S T U V W X Y Z A B C D E F G H I J K L M N O P Q 18 S T U V W X Y Z A B C D E F G H I J K L M N O P Q R 19 T U V W X Y Z A B C D E F G H I J K L M N O P Q R S 20 U V W X Y Z A B C D E F G H I J K L M N O P Q R S T 21 V W X Y Z A B C D E F G H I J K L M N O P Q R S T U 22 W X Y Z A B C D E F G H I J K L M N O P Q R S T U V 23 X Y Z A B C D E F G H I J K L M N O P Q R S T U V W 24 Y Z A B C D E F G H I J K L M N O P Q R S T U V W X 25 Z A B C D E F G H I J K L M N O P Q R S T U V W X Y Table 1.4 Polyalphabetic substition table
Chapter 2 Stream Ciphers If we have a more detailed look at the types of cryptographic algorithms that exist, we see that symmetric ciphers can be divided into two families, stream ciphers and block ciphers, as shown in Figure 2.1. Fig. 2.1 Main areas within cryptography This chapter is concerned with stream ciphers, which are an important class of cryptographic primitives. In the chapter you will learn: The pros and cons of stream ciphers Random and pseudorandom number generators A truly unbreakable cipher: the one-time pad (OTP) Linear feedback shift registers The modern stream ciphers ChaCha20, Salsa20 and Trivium 37 C. Paar et al., Understanding Cryptography, https://doi.org/10.1007/978-3-662-69007-9_2 © The Editor(s) (if applicable) and The Author(s), under exclusive license to Springer-Verlag GmbH, DE, part of Springer Nature 2024
38 2 Stream Ciphers 2.1 Introduction This section will first discuss the difference between stream and block ciphers, and then introduce the principle way stream ciphers work. 2.1.1 Stream Ciphers vs. Block Ciphers Symmetric cryptography is split into block ciphers and stream ciphers, which are easy to distinguish. Figure 2.2 depicts the operational differences between stream (Figure 2.2a) and block ciphers (Figure 2.2b). In both cases, we want to encrypt b bits at a time, where b is the width of the block cipher. A description of the operation of the two types of symmetric ciphers follows. (a) (b) Fig. 2.2 Principles of encrypting b bits with a stream (a) and a block (b) cipher Stream ciphers encrypt bits individually. This is achieved by adding a bit from a key stream to a plaintext bit. There are synchronous stream ciphers where the key
2.1 Introduction 39 stream depends only on the key, and asynchronous ones where the key stream also depends on the ciphertext. If the dotted line in Figure 2.3 is present, the stream cipher is an asynchronous one. Most practical stream ciphers are synchronous ones and Section 2.4 of this chapter will introduce three modern ciphers of this type. An example of an asynchronous stream cipher is the cipher feedback (CFB) mode, introduced in Section 5.1.4. Fig. 2.3 Synchronous and asynchronous stream ciphers Block ciphers encrypt a block of b plaintext bits at a time with the same key. The ciphers are constructed such that the encryption of any plaintext bit in a given block depends on every other plaintext bit in the same block. In practice, the vast majority of block ciphers either have a block length of 128 bits (16 bytes) like the Advanced Encryption Standard (AES), or a block length of 64 bits (8 bytes) like the Data Encryption Standard (DES) or the triple DES (3DES). These two ciphers are introduced in Chapters 4 and 3, respectively. In practice, in particular for encrypting computer communication on the inter- net, block ciphers are used more often than stream ciphers. In the earlier days of modern cryptography, roughly in the 1980s and 1990s, it was assumed that stream ciphers could encrypt more efficiently than block ciphers. They were particularly relevant for applications with low computational resources, e.g., for cell phones or other small embedded devices. Back then, stream ciphers were often realized in hardware. A prominent example of such an algorithm is the A5/1 cipher, which is used for voice encryption in the (soon to be outdated) GSM mobile phone standard. Since then, many block ciphers that are specifically designed with hardware effi- ciency in mind have been proposed, such as the PRESENT algorithms, which is described in Section 3.7.3. At the same time, there are modern stream ciphers that are very well suited for high-speed software implementations, including Salsa20 (cf. Section 2.4.1) and ChaCha (cf. Section 2.4.2).
40 2 Stream Ciphers 2.1.2 Encryption and Decryption with Stream Ciphers As mentioned above, stream ciphers encrypt plaintext bits individually. The question now is: How does encryption of an individual bit work? The answer is surprisingly simple: Each bit xi is encrypted by adding a secret key stream bit si modulo 2. Definition 2.1.1 Stream Cipher Encryption and Decryption The plaintext, the ciphertext and the key stream consist of individ- ual bits, i.e., xi, yi, si ∈ {0, 1}. Encryption: yi = esi (xi) ≡ xi + si mod 2 Decryption: xi = dsi (yi) ≡ yi + si mod 2 Since the encryption and decryption functions are both simple additions modulo 2, we can depict the basic operation of a stream cipher as shown in Figure 2.4. Note that we use a circle with an addition sign as the symbol for modulo 2 addition. Fig. 2.4 Encryption and decryption with stream ciphers Just looking at the formulae in the definition, there are three observations about the stream cipher encryption and decryption function which we should clarify: 1. Why are encryption and decryption the same function? 2. Why can we use a simple modulo 2 addition as encryption? 3. What is the nature of the key stream bits si? The following discussion of these three items will already give us an understanding of some important stream cipher properties. Why Are Encryption and Decryption the Same Function? The reason for the similarity of the encryption and decryption functions can easily be shown. We perform what is called a proof of correctness, i.e., we show that the decryption function actually produces the plaintext bit xi again. We know that
2.1 Introduction 41 ciphertext bit yi was computed using the encryption function yi ≡ xi + si mod 2. We insert this encryption expression in the decryption function: dsi (yi) ≡ yi + si mod 2 ≡ (xi + si) + si mod 2 ≡ xi + si + si mod 2 ≡ xi + 2 si mod 2 ≡ xi + 0 mod 2 ≡ xi mod 2 Q.E.D. The trick here is that the expression (2 si mod 2) always has the value zero since 2 ≡ 0 mod 2. Another way of understanding this is as follows: If si has the value 0, then 2 si = 2 · 0 ≡ 0 mod 2. If si = 1, we have 2 si = 2 · 1 = 2 ≡ 0 mod 2. Why is Modulo 2 Addition a Good Encryption Function? A mathematical explanation for this is given in the context of the one-time pad in Section 2.2.2. However, it is worth having a closer look at addition modulo 2. If we do arithmetic modulo 2, the only possible values are 0 and 1 (because if you divide by 2 only two remainders can occur, 0 and 1). Thus, we can treat arithmetic modulo 2 as Boolean functions such as logic AND, OR, NAND, etc. Let’s look at the truth table for modulo 2 addition: xi si yi ≡ xi + si mod 2 0 0 0 0 1 1 1 0 1 1 1 0 This should look familiar to most readers: It is the truth table of the exclusive-OR, also called XOR, gate. This is an important fact: Modulo 2 addition is equivalent to the XOR operation. The XOR operation plays a major role in modern cryptography and will be used many times in the remainder of this book. The question now is, why is the XOR operation so useful, as opposed to, say, the AND operation? Let’s assume we want to encrypt the plaintext bit xi = 0. If we look at the truth table we find that we are on either the 1st or 2nd line: xi si y i 0 0 0 0 1 1 1 0 1 1 1 0
42 2 Stream Ciphers Depending on the key bit, the ciphertext yi is either a zero (si = 0) or one (si = 1). If the key bit si behaves perfectly randomly, i.e., it is unpredictable and has exactly a 50% chance to have each of the values 0 and 1, then both possible ciphertexts also occur with a 50% likelihood. Likewise, if we encrypt the plaintext bit xi = 1, we are on line 3 or 4 of the truth table. Again, depending on the value of the key stream bit si, there is a 50% chance that the ciphertext is each of the values 0 and 1. We just observed that the XOR function is perfectly balanced, i.e., by observing an output value, there is exactly a 50% chance for any value of the input bits. This distinguishes the XOR gate from other Boolean functions such as the OR, AND or NAND gates. Moreover, AND and NAND gates are not invertible. Let’s look at a very simple example of encryption with a stream cipher. Example 2.1. Alice wants to encrypt the letter A, where the letter is given in ASCII code. The ASCII value for A is 6510 = 10000012. Let’s furthermore assume that the first key stream bits are (s0, . . . , s6) = 0101100. Alice Oscar Bob x0, . . . , x6 = 1000001 = A ⊕ s0, . . . , s6 = 0101100 y0, . . . , y6 = 1101101 = m m=1101101 −−−−−−−−−−−−→ y0, . . . , y6 = 1101101 ⊕ s0, . . . , s6 = 0101100 x0, . . . , x6 = 1000001 = A Note that encryption by Alice turns the uppercase A into the lowercase letter m. Oscar, the attacker who eavesdrops on the channel, only sees the ciphertext letter m. Decryption by Bob with the same key stream reproduces the plaintext A again. So far, stream ciphers look unbelievably easy: One simply takes the plaintext, performs an XOR operation with the key and obtains the ciphertext. On the receiving side, Bob does the same. The “only” thing left to discuss is the last question from above. What Exactly is the Nature of the Key Stream? It turns out that the generation of the values si, which are called the key stream, is the central issue for the security of stream ciphers. In fact, the security of a stream cipher completely depends on the key stream. The key stream bits si are not the key bits themselves. So, how do we get the key stream? Generating the key stream is pretty much what stream ciphers are about. This is a major topic and is discussed later in this chapter. However, we can already guess that a central requirement for the key stream bits should be that they appear like a random sequence to an attacker. Otherwise, an attacker Oscar could guess the bits and do the decryption himself. Hence, we first need to learn more about random numbers.
2.2 Random Numbers and an Unbreakable Stream Cipher 43 2.2 Random Numbers and an Unbreakable Stream Cipher 2.2.1 Random Number Generators As we saw in the previous section, the actual encryption and decryption of stream ciphers is extremely simple. The security of stream ciphers hinges entirely on a “suitable” key stream s0, s1, s2, . . . Since randomness plays a major role, we will first learn about the three types of random number generators (RNG) that are important for us. True Random Number Generators (TRNGs) True random number generators (TRNGs) are characterized by the fact that their output cannot be reproduced. For instance, if we flip a coin 100 times and record the resulting sequence of 100 bits, it will be virtually impossible for anyone on Earth to generate the same 100-bit sequence. The chance of success is 1/2100, which is an ex- tremely small probability. Ideally, TRNGs are based on physical processes that can- not be reproduced. Examples include coin flipping or rolling of dice by humans. In computer systems, modern CPUs are often equipped with hardware-based TRNGs or else there is a TPM (trusted platform module) on the motherboard which contains a TRNG. In computer systems without a hardware TRNG, random processes within the computer are used as entropy sources, e.g., fine-grained timing measurements of interrupts or other random data from device drivers. In cryptography, TRNGs are often needed for generating session keys, which are then distributed between Alice and Bob, but also for other purposes such as generation of nonces. (General) Pseudorandom Number Generators (PRNGs) Pseudorandom number generators (PRNGs) generate sequences which are com- puted from an initial seed value. Often they are computed recursively in the follow- ing way: s0 = seed si+1 = f (si), i = 0, 1, . . . A generalization of this is generators of the form si+1 = f (si, si−1, . . . , si−t ), where t is a fixed integer. A popular example is the linear congruential generator: s0 = seed si+1 ≡ a si + b mod m, i = 0, 1, . . .
44 2 Stream Ciphers where a, b, m are integer constants. Note that PRNGs are not random in a true sense because they can be computed and are thus completely deterministic. A widely used example is the rand() function used in ANSI C. It has the parameters: s0 = 12345 si+1 ≡ 1103515245 si + 12345 mod 231, i = 0, 1, . . . A common requirement of PRNGs is that they possess good statistical proper- ties, meaning their output approximates a sequence of true random numbers. There are many mathematical tests, e.g., the chi-square test, which can verify the statistical behavior of PRNG sequences. Note that there are many, many applications for pseu- dorandom numbers outside cryptography. For instance, many types of simulations or testing, e.g., of software or of VLSI chips, need random test data as input. This is the reason why a PRNG is included in the ANSI C specification. Cryptographically Secure Pseudorandom Number Generators (CSPRNGs) Cryptographically secure pseudorandom number generators (CSPRNGs) are a spe- cial type of PRNG that possess the following additional property: A CSPRNG is a PRNG which is unpredictable. Informally, this means that given n output bits of the key stream si, si+1, . . . , si+n−1, where n is some integer, it is computationally infea- sible to compute the subsequent bits si+n, si+n+1, . . . A more exact definition is that given n consecutive bits of the key stream, there is no polynomial-time algorithm that can predict the next bit sn+1 with better than 50% chance of success. Another property of CSPRNGs is that given the above sequence, it should be computation- ally infeasible to compute any preceding bits si−1, si−2, . . . Note that the need for unpredictability of CSPRNGs is unique to cryptography. In virtually all other situations where pseudorandom numbers are needed in computer science or engineering, unpredictability is not needed. As a consequence, the dis- tinction between PRNGs and CSPRNGs and its relevance for stream ciphers is often not clear to non-cryptographers. Almost all PRNGs that were designed without the clear purpose of being stream ciphers are not CSPRNGs. 2.2.2 The One-Time Pad In the following we discuss what happens if we use the three types of random num- bers as generators for the key stream sequence s0, s1, s2, . . . of a stream cipher. Let’s first define what a perfect cipher should be.
2.2 Random Numbers and an Unbreakable Stream Cipher 45 Definition 2.2.1 Unconditional Security A cryptosystem is unconditionally or information-theoretically se- cure if it cannot be broken even with infinite computational re- sources. Unconditional security is based on information theory and assumes no limit on the attacker’s computational power. This looks like a pretty straightforward defini- tion. It is in fact straightforward, but the requirements for a cipher to be uncondition- ally secure are tremendous. Let’s look at it using a gedankenexperiment: Assume we have a symmetric encryption algorithm (it doesn’t matter whether it’s a block cipher or stream cipher) with a key length of 10,000 bits, and the only attack that works is an exhaustive key search, i.e, a brute-force attack. From the discussion in Sec- tion 1.3.2, we recall that 128 bits are more than enough for long-term security. So, is a cipher with 10,000 bits unconditionally secure? The answer is a resounding: No! Since an attacker can have infinite computational resources, we can simply assume that the attacker has 210000 computers available and every computer checks exactly one key. This will give us a correct key in one time step. Of course, there is no way that 210000 computers can ever be built, the number is too large. (It is estimated that there are “only” about 2266 atoms in the known universe.) The cipher would merely be computationally secure, but not unconditionally. All this said, we now show a way to build an unconditionally secure cipher that is quite simple. This cipher is called the one-time pad. Definition 2.2.2 One-Time Pad (OTP) A stream cipher for which 1. the key stream s0, s1, s2, . . . is generated by a true random num- ber generator, and 2. the key stream is only known to the legitimate communicating parties, and 3. every key stream bit si is only used once is called a one-time pad. The one-time pad is unconditionally se- cure. It is easy to show why the OTP is unconditionally secure. Here is a sketch of a proof. For every ciphertext bit we get an equation of the form: y0 ≡ x0 + s0 mod 2 y1 ≡ x1 + s1 mod 2 ... Each individual relation is a linear equation modulo 2 with two unknowns. It is impossible to derive a unique solution for such equations. If the attacker knows the
46 2 Stream Ciphers value for y0 (0 or 1), he cannot determine the value of x0. In fact, the solutions x0 = 0 and x0 = 1 are exactly equally likely if s0 stems from a truly random source and there is 50% chance that it has each of the values 0 and 1. The situation is identical for the second equation and all subsequent ones. Note that the situation is different if the values si are not truly random. In this case, there is some functional relationship between them, and the equations shown above are not independent. Even though it might still be hard to solve the system of equations, it is not provably secure! Great, now we have a simple cipher that is perfectly secure. There are rumors that the red telephone between the White House and the Kremlin was encrypted using an OTP during the Cold War. Obviously there must be a catch since OTPs are not used for web browsers, email encryption, instant messaging on smartphones, or other im- portant modern applications. Let’s look at the implications of the three requirements in Definition 2.2.2. The first requirement means that we need a TRNG. For this we need a device, e.g., based on white noise of a semiconductor, that generates truly random bits. Since standard PCs often have TRNGs nowadays, this requirement can be met without too much effort. The second requirement means that Alice has to get the random bits securely to Bob. In practice that could mean that Alice stores the true random bits on an USB stick or a portable SSD and sends them securely, e.g., with a trusted courier, to Bob. This is not great but doable in certain applications. The third requirement is the most impractical one: Key stream bits cannot be re- used. This implies that we need one key bit for every bit of plaintext. Hence, our key is as long as the plaintext! This is probably the major drawback of the OTP. Even if Alice and Bob share 1 GByte of true random numbers, we run quickly into limits. Just exchanging a few large files could exhaust the 1 GByte of key material. After that they would need to repeat the cumbersome and slow process of exchanging true random key stream bits again, e.g., by a trusted courier. For these reasons OTPs are rarely used in practice. However, they give us a great design idea for secure ciphers: If we XOR truly random bits and plaintext, we get ciphertext that can certainly not be broken by an attacker. We will see in the next section how we can use this fact to build practical stream ciphers. 2.2.3 Towards Practical Stream Ciphers In the previous section we saw that OTPs are unconditionally secure, but they have drawbacks which make them impractical. What we try to do with practical stream ciphers is to replace the truly random key stream bits with a pseudorandom number generator where the key k serves as a seed. The principle of practical stream ciphers is shown in Figure 2.5. Before we turn to stream ciphers used in the real world, it should be stressed that practical stream ciphers are not unconditionally secure. In fact, all known practi- cal cryptographic algorithms (stream ciphers, block ciphers, public-key algorithms) are not unconditionally secure. The best we can aim for is computational security, which we define as follows.
2.2 Random Numbers and an Unbreakable Stream Cipher 47 Fig. 2.5 Practical stream ciphers Definition 2.2.3 Computational Security A cryptosystem is computationally secure if the best known algo- rithm for breaking it requires at least t operations. This seems like a reasonable definition but there are still several problems with it. First, often we do not know what the best algorithm for a given attack is. A prime example is the RSA public-key scheme, which can be broken by factoring large in- tegers. Even though many factoring algorithms are known, we do not know whether there exist any better ones. Second, even if a lower bound on the complexity of one attack is known, we do not know whether any other, more powerful attacks are possible. We saw this in Section 1.2.2 during the discussion about the substitution cipher: Even though we know the exact computational complexity for an exhaustive key search, there exist other, more powerful attacks. The best we can do in practice is to design cryptographic schemes for which it is assumed that they are computa- tionally secure. For symmetric ciphers this usually means one hopes that there is no attack method with a complexity better than an exhaustive key search. Let’s go back to Figure 2.5. This design emulates (“behaves to a certain extent like”) a one-time pad. It has the major advantage over the OTP that Alice and Bob only need to exchange a secret key that is at most a few 100 bits long, and that does not have to be as long as the message we want to encrypt. We now have to think carefully about the properties of the key stream s0, s1, s2, . . . that is generated by Alice and Bob. Obviously, we need some type of random number generator to derive the key stream. First, we note that we cannot use a TRNG since, by definition, Alice and Bob will not be able to generate the same key stream. Instead we need deterministic, i.e., pseudorandom, number generators. We now look at the other two generators that were introduced in the previous section.
48 2 Stream Ciphers Building Key Streams from PRNGs Here is an idea that seems promising (but in fact is pretty bad): Many PRNGs pos- sess good statistical properties, which are necessary for a strong stream cipher. If we apply statistical tests to the key stream sequence, the output should pretty much behave like the bit sequence generated by tossing a coin. So it is tempting to assume that a PRNG can be used to generate the key stream. But all of this is not sufficient for a stream cipher since our opponent, Oscar, is smart. Consider the following at- tack. Example 2.2. Let’s assume a PRNG based on the linear congruential generator: S0 = seed Si+1 ≡ A Si + B mod m, i = 0, 1, . . . where we choose m to be 100 bits long and Si, A, B ∈ {0, 1, . . . , m − 1}. Note that this PRNG can have excellent statistical properties if we choose the parameters carefully. The modulus m is part of the encryption scheme and is publicly known. The secret key comprises the values (A, B), each with a length of 100 bits. That gives us a key length of 200 bits, which is more than sufficient to protect against a brute-force attack. Since this is a stream cipher, Alice can encrypt: yi ≡ xi + si mod 2 where si are the bits of the binary representation of the PRNG output symbols S j. But Oscar can easily launch an attack. Assume he knows the first 300 bits of plaintext (this is only 300/8=37.5 bytes), e.g., file header information or he guesses part of the plaintext. Since he certainly knows the ciphertext, he can now compute the first 300 bits of key stream as: si ≡ yi + xi mod m , i = 1, 2, . . . , 300 These 300 bits immediately give the first three output symbols of the PRNG: S1 = (s1, . . . , s100), S2 = (s101, . . . , s200) and S3 = (s201, . . . , s300). Oscar can now generate two equations: S2 ≡ A S1 + B mod m S3 ≡ A S2 + B mod m This is a system of linear equations over Zm with two unknowns A and B. But those two values are the key, and we can immediately solve the system, yielding: A ≡ (S2 − S3)/(S1 − S2) mod m B ≡ S2 − S1(S2 − S3)/(S1 − S2) mod m In case gcd((S1 − S2), m)) 6 = 1 we get multiple solutions since this is an equation system over Zm. However, with a fourth piece of known plaintext the key can be
2.3 Shift Register-Based Stream Ciphers 49 uniquely detected in almost all cases. Alternatively, Oscar simply tries to encrypt the message with each of the multiple solutions found. Hence, in summary: If we know a few pieces of plaintext, we can compute the key and decrypt the entire ciphertext! This type of attack is why the notion of CSPRNG was invented. Building Key Streams from CSPRNGs What we need to do to prevent the attack above is to use a CSPRNG, which ensures that the key stream is unpredictable. We recall that this means that given the first n output bits of the key stream s1, s2, . . . , sn, it is computationally infeasible to com- pute the bits sn+1, sn+2, . . . Unfortunately, pretty much all pseudorandom number generators that are used for applications outside cryptography are not cryptograph- ically secure. Hence, in practice, we need to use pseudorandom number generators that are specially designed for stream ciphers. The question now is how practical stream ciphers actually look. There are many proposals for stream ciphers in the literature. They can roughly be classified as ci- phers either optimized for software implementation or optimized for hardware im- plementation. In the former case, the ciphers typically require few CPU instructions to compute one key stream bit. In the latter case, they tend to be based on opera- tions that can easily be realized in hardware. A popular example is shift registers with feedback, which are discussed in the next section. Another class of stream ci- phers is built upon pseudorandom functions based on 32-bit addition, rotation and XOR operations, so called add-rotate-XOR operations, and will be discussed in Sec- tion 2.4. A third class of stream ciphers uses block ciphers as building blocks. The cipher feedback mode, output feedback mode and counter mode to be introduced in Chapter 5 are examples of stream ciphers derived from block ciphers. 2.3 Shift Register-Based Stream Ciphers As we have learned so far, practical stream ciphers use a stream of key bits s1, s2, . . . that are generated by the key stream generator, which should have certain properties. An elegant way of realizing long pseudorandom sequences is to use linear feedback shift registers (LFSRs). They are easily implemented in hardware and many, but certainly not all, stream ciphers make use of LFSRs. A prominent example is the A5/1 cipher, which is standardized for voice encryption in the (somewhat outdated) GSM mobile communication standard. As we will see, even though a plain LFSR produces a sequence with good statistical properties, it is cryptographically weak. However, combinations of LFSRs can make secure stream ciphers. Trivium, intro- duced in Section 2.4.3, is an example of such a cipher. It should be stressed that there are many ways to construct stream ciphers, as we will see in Section 2.4.
50 2 Stream Ciphers 2.3.1 Linear Feedback Shift Registers (LFSRs) An LFSR consists of clocked storage elements (flip-flops) and a feedback path. The number of storage elements gives us the degree of the LFSR. In other words, an LFSR with m flip-flops is said to be of degree m. The feedback network computes the input for the last flip-flop as the XOR-sum of certain flip-flops in the shift regis- ter. Example 2.3. Simple LFSR We consider an LFSR of degree m = 3 with flip-flops FF2, FF1, FF0, and a feedback path as shown in Figure 2.6. The internal state bits are denoted by si and are shifted by one to the right with each clock tick. The rightmost state bit is also the current output bit. The leftmost state bit is computed in the feedback path, which is the XOR sum of some of the flip-flop values in the previous clock period. Since the XOR is a linear operation, such circuits are called linear feedback shift registers. If we assume an initial state of (s2 = 1, s1 = 0, s0 = 0),CLK FF FF FF Fig. 2.6 Linear feedback shift register of degree 3 with initial values s2, s1, s0 Table 2.1 gives the complete sequence of states of the LFSR. Note that the rightmost Table 2.1 Sequence of states of the LFSR clk FF2 FF1 FF0 = si 0 1 0 0 1 0 1 0 2 1 0 1 3 1 1 0 4 1 1 1 5 0 1 1 6 0 0 1 7 1 0 0 8 0 1 0
2.3 Shift Register-Based Stream Ciphers 51 column is the output of the LFSR. One can see from this example that the LFSR starts to repeat after clock cycle 6. This means the LFSR output has period of length 7 and has the form: 0010111 0010111 0010111 . . . There is a simple formula which determines the functioning of this LFSR. Let’s look at how the output bits si are computed, assuming the initial state bits s0, s1, s2: s3 ≡ s1 + s0 mod 2 s4 ≡ s2 + s1 mod 2 s5 ≡ s3 + s2 mod 2 ... In general, the output bit is computed as: si+3 ≡ si+1 + si mod 2 where i = 0, 1, 2, . . . This was, of course, a simple example. However, we could already observe many important properties. We will now look at general LFSRs. A Mathematical Description of LFSRs The general form of an LFSR of degree m is shown in Figure 2.7. It shows m flip-flops and m possible feedback locations, all combined by the XOR operation. Whether a feedback path is active or not is defined by the feedback coefficients pm−1, . . . , p0, p1, which have the following function: If pi = 1 (closed switch), the feedback is active. If pi = 0 (open switch), the corresponding flip-flop output is not used for the feedback. With this notation, we obtain an elegant mathematical description for the feedback path. If we multiply the output of flip-flop i by its coefficient pi, the result is either the output value if pi = 1, which corresponds to a closed switch, or the value zero if pi = 0, which corresponds to an open switch. The values of the feedback coefficients are crucial for the output sequence produced by the LFSR. Let’s assume the LFSR is initially loaded with the values sm−1, . . . , s0. The next output bit of the LFSR sm, which is also the input to the leftmost flip-flop, can be computed by the XOR-sum of the products of flip-flop outputs and corresponding feedback coefficients: sm ≡ sm−1 pm−1 + · · · + s1 p1 + s0 p0 mod 2
52 2 Stream CiphersCLK FF FF FF Fig. 2.7 General LFSR with feedback coefficients pi and initial values sm−1, . . . , s0 The next LFSR output can be computed as: sm+1 ≡ sm pm−1 + · · · + s2 p1 + s1 p0 mod 2 In general, the output sequence can be described as: sm+i ≡ m−1 ∑ j=0 p j · si+ j mod 2; si, p j ∈ {0, 1}; i = 0, 1, 2, . . . (2.1) Clearly, the output values are given through a combination of some previous output values. LFSRs are sometimes referred to as linear recurrences. Due to the finite number of recurring states, the output sequence of an LFSR repeats periodically. This was also illustrated in Example 2.3, where the period was 7. Moreover, an LFSR can produce output sequences of different lengths, depending on the feedback coefficients. The following theorem gives us the maximum length of an LFSR as function of its degree. Theorem 2.3.1 The maximum sequence length generated by an LFSR of degree m is 2m − 1. It is easy to show that this theorem holds. The state of an LFSR is uniquely deter- mined by the m internal register bits. Given a certain state, the LFSR deterministi- cally assumes its next state. Because of this, as soon as an LFSR assumes a previous state, it starts to repeat. Since an m-bit state vector can only assume 2m − 1 nonzero states, the maximum sequence length before repetition is 2m − 1. Note that the all- zero state must be excluded. If an LFSR assumes this state, it will get “stuck” in it, i.e., it will never be able to leave it again. Note that only certain configurations (p0, . . . , pm−1) yield maximum-length LFSRs. We give a small example of this be- low.
2.3 Shift Register-Based Stream Ciphers 53 Example 2.4. LFSR with maximum-length output sequence Given an LFSR of degree m = 4 and the feedback path (p3 = 0, p2 = 0, p1 = 1, p0 = 1), the output sequence of the LFSR has a period of 2m − 1 = 15, i.e., it is a maximum-length LFSR. Example 2.5. LFSR with non-maximum output sequence Given an LFSR of degree m = 4 and (p3 = 1, p2 = 1, p1 = 1, p0 = 1), then the output sequence has period of 5; therefore, it is not a maximum-length LFSR. The mathematical background of the properties of LFSR sequences is beyond the scope of this book. However, we conclude this introduction to LFSRs with some additional facts. LFSRs are often specified by polynomials using the following no- tation: An LFSR with a feedback coefficient vector (pm−1, . . . , p1, p0) is represented by the polynomial: P(x) = xm + pm−1xm−1 + . . . + p1x + p0 For instance, the LFSR from the example above with coefficients (p3 = 0, p2 = 0, p1 = 1, p0 = 1) can alternatively be specified by the polynomial x4 + x + 1. This seemingly odd notation as a polynomial has several advantages. For instance, maximum-length LFSRs have what is called primitive polynomials. Primitive poly- nomials are a special type of irreducible polynomial. Irreducible polynomials are roughly comparable with prime numbers, i.e., their only factors are 1 and the polynomial itself. Primitive polynomials can relatively easily be computed. Hence, maximum-length LFSRs can easily be found. Table 2.2 shows one primitive poly- nomial for every value of m in the range from m = 2, 3, . . . , 128. As an example, the notation (0, 2, 5) refers to the polynomial 1 + x2 + x5. Note that there are many primitive polynomials for every given degree m. For instance, there exist 69,273,666 different primitive polynomials of degree m = 31. 2.3.2 Known-Plaintext Attack Against Single LFSRs As indicated by its name, LFSRs are linear. Linear systems are governed by linear relationships between their inputs and outputs. Since linear dependencies can rela- tively easily be analyzed, this can be a major advantage in many application areas, e.g., in communication systems. However, a cryptosystem where the key bits only occur in linear relationships makes a highly insecure cipher. We will now investigate how the linear behavior of an LFSR leads to a powerful attack. If we use an LFSR as a stream cipher, the secret key k is the feedback coefficient vector (pm−1, . . . , p1, p0). An attack is possible if the attacker Oscar knows some plaintext and the corresponding ciphertext. We further assume that Oscar knows the degree m of the LFSR. The attack is so efficient that he could also easily try a large number of possible m values, so that this assumption is not a major restriction. Let the known plaintext be given by x0, x1, . . . , x2m−1 and the corresponding ciphertext
54 2 Stream Ciphers Table 2.2 Primitive polynomials for maximum-length LFSRs (0,1,2) (0,1,3,4,24) (0,1,2,3,5,8,46) (0,1,5,7,68) (0,2,3,5,90) (0,1,2,3,6,8,112) (0,1,3) (0,3,25) (0,5,47) (0,2,5,6,69) (0,1,2,3,5,6,7,91) (0,2,3,5,113) (0,1,4) (0,1,2,6,26) (0,1,2,4,5,7,48) (0,1,3,5,70) (0,2,5,6,92) (0,2,3,6,7,8,114) (0,2,5) (0,1,2,5,27) (0,4,5,6,49) (0,1,3,5,71) (0,2,93) (0,1,2,3,5,7,115) (0,1,6) (0,3,28) (0,2,3,4,50) (0,1,2,3,4,6,72) (0,1,5,6,94) (0,2,5,6,116) (0,1,7) (0,2,29) (0,1,3,6,51) (0,2,3,4,73) (0,1,2,4,5,6,95) (0,1,2,5,117) (0,2,3,4,8) (0,1,4,6,30) (0,3,52) (0,3,4,7,74) (0,2,3,4,6,7,96) (0,2,5,6,118) (0,4,9) (0,3,31) (0,1,2,6,53) (0,1,3,6,75) (0,6,97) (0,8,119) (0,3,10) (0,1,2,3,5,7,32) (0,2,3,4,5,6,54) (0,2,4,5,76) (0,1,2,3,4,7,98) (0,1,2,5,6,7,120) (0,2,11) (0,1,4,6,33) (0,1,2,6,55) (0,2,5,6,77) (0,4,5,7,99) (0,1,5,8,121) (0,1,4,6,12) (0,1,2,5,6,7,34) (0,2,4,7,56) (0,1,2,7,78) (0,2,7,8,100) (0,1,2,6,122) (0,1,3,4,13) (0,2,35) (0,2,3,5,57) (0,2,3,4,79) (0,1,6,7,101) (0,2,123) (0,1,3,5,14) (0,1,2,4,5,6,36) (0,1,5,6,58) (0,1,2,3,5,7,80) (0,3,5,6,102) (0,5,6,7,124) (0,1,15) (0,1,2,3,4,5,37) (0,1,3,4,5,6,59) (0,4,81) (0,2,3,4,5,7,103) (0,1,2,3,5,7,125) (0,2,3,5,16) (0,1,5,6,38) (0,1,60) (0,1,4,6,7,8,82) (0,2,3,4,5,6,8,9,104) (0,2,4,7,126) (0,3,17) (0,4,39) (0,1,2,5,61) (0,2,4,7,83) (0,1,2,4,5,6,105) (0,1,127) (0,1,2,5,18) (0,3,4,5,40) (0,3,5,6,62) (0,1,3,5,7,8,84) (0,1,5,6,106) (0,1,2,7,128) (0,1,2,5,19) (0,3,41) (0,1,63) (0,1,2,8,85) (0,1,2,3,5,7,107) (0,3,20) (0,1,2,3,4,5,42) (0,1,3,4,64) (0,2,5,6,86) (0,1,2,3,4,5,6,7,9,10,108) (0,2,21) (0,3,4,6,43) (0,1,3,4,65) (0,1,5,7,87) (0,2,4,5,109) (0,1,22) (0,2,5,6,44) (0,2,3,5,6,8,66) (0,1,3,4,5,8,88) (0,1,4,6,110) (0,5,23) (0,1,3,4,45) (0,1,2,5,67) (0,3,5,6,89) (0,2,4,7,111) by y0, y1, . . . , y2m−1. With these 2m pairs of plaintext and ciphertext bits, Oscar re- constructs the first 2m key stream bits: si ≡ xi + yi mod 2; i = 0, 1, . . . , 2m − 1. The goal now is to find the key, i.e., the m feedback coefficients pi. Equation (2.1) is a description of the relationship of the unknown key bits pi and the key stream output. We repeat the equation here for convenience: sm+i ≡ m−1 ∑ j=0 p j · si+ j mod 2; si, p j ∈ {0, 1}; i = 0, 1, 2, . . . Note that we get a different equation for every value of i. Moreover, the equations are linearly independent. With this knowledge, Oscar can generate m equations for the first m values of i: i = 0, sm ≡ pm−1sm−1 + . . . + p1s1 + p0s0 mod 2 i = 1, sm+1 ≡ pm−1sm + . . . + p1s2 + p0s1 mod 2 ... ... ... ... ... i = m − 1, s2m−1 ≡ pm−1s2m−2 + . . . + p1sm + p0sm−1 mod 2 (2.2) He now has m linear equations with m unknowns p0, p1, . . . , pm−1. This system can easily be solved by Oscar using Gaussian elimination, matrix inversion or any other algorithm for solving systems of linear equations. Even for large values of m, this can be done easily with a standard PC. This situation has major consequences: as soon as Oscar knows 222mmm output bits of an LFSR of degree mmm, the pppiii coefficients can be exactly constructed by merely
2.4 Practical Stream Ciphers 55 solving a system of linear equations. Once he has computed these feedback coef- ficients, he can “build” the LFSR and load it with any m consecutive output bits that he already knows. Oscar can now clock the LFSR and produce the entire output sequence. Because of this powerful attack, LFSRs by themselves are extremely inse- cure! They are a good example of a PRNG with good statistical properties but with terrible cryptographical ones. Nevertheless, all is not lost. There are many stream ciphers which use combinations of several LFSRs to build strong cryptosystems. The cipher Trivium in Section 2.4.3 is an example. 2.4 Practical Stream Ciphers Even though stream ciphers had been popular in the “early days” of modern cryptog- raphy, roughly in the 1980s, block ciphers became more dominant during the 1990s. This development was in part due to the successful attacks against many of the early stream ciphers. This was a main motivation for the eSTREAM project, which was organized by a network of European cryptographers. In 2004, eSTREAM issued a call for new stream ciphers and the selection process ended in 2008. The ciphers were divided into two “profiles”. Profile 1 contains stream ciphers that allow high- throughput software implementations. Profile 2 algorithms are hardware-friendly stream ciphers, which have a low gate count and power consumption. In this section we will introduce representatives for each of the two profiles: Salsa20 together with its variant ChaCha (Profile 1) and Trivium (Profile 2). 2.4.1 Salsa20 Salsa20 is a family of software-efficient stream ciphers developed by Daniel J. Bern- stein in 2005. The cipher uses a pseudorandom function based on 32-bit additions, rotations and XOR operations. Such algorithms are referred to as add-rotate-XOR (ARX) ciphers. The original cipher has 20 rounds and is denoted by Salsa20/20. This cipher is already faster than AES on most CPUs. Subsequently, Bernstein in- troduced two Salsa20 variants with a reduced round count, named Salsa20/12 and Salsa20/8, which are even faster. No attacks are known against any of the Salsa20 variants that are better than a brute-force attack and the cipher is considered to be very secure. In the following we will describe Salsa20 with 20 rounds. Encryption and Decryption with Salsa20 Salsa20 supports key lengths of 256 and 128 bits. However, the designer recom- mends 256 bits. The core of Salsa20 is a function with a 512-bit input and a 512- bit output. For both encryption and decryption, Salsa20 processes the key, a nonce
56 2 Stream Ciphers (which stands for “number used only once”) and a 64-bit block number, and gen- erates a 512-bit block of key stream. Like every stream cipher, the result of this process is XOR-added with the plaintext (encryption) or ciphertext (decryption). Since Salsa20 generates an output of 512 bits, one can encrypt (or decrypt) 512 plaintext or ciphertext bits at once. The main purpose of the nonce is that two key streams produced by the cipher should be different, even though the key has not changed. If this were not the case, the following attack becomes possible: If an attacker has a known plaintext from the first encryption, he can compute the corresponding key stream. The second encryption using the same key stream can now immediately be deciphered. With- out a changing nonce, stream cipher encryption is highly deterministic. Nonces are also often used for IVs (initialization vectors), which are used in many cipher con- structions. Methods for generating IVs are discussed in Section 5.1.2. We note that nonces and IVs do not have to be kept secret; It must merely be ensured that nonces change for every session. Encryption and decryption can be expressed as follows. Let k be a 32-byte (256 bits) or 16-byte (128 bits) sequence. Let n be an 8-byte nonce. Let x be an l-byte message for some l ∈ {0, 1, . . . , 270}. The Salsa20 encryption of a message x yields an l-byte ciphertext y: y = Salsa20k(n) ⊕ x For decryption, the same key stream is used and XORed to the ciphertext, i.e., x = Salsa20k(n) ⊕ y Since each block depends only on the key, the nonce and the block number, the key stream blocks can be computed independently of each other and blocks can be computed in parallel. This is advantageous for high-speed implementations of Salsa20. Core Function of Salsa20 The 512-bit internal state of Salsa20 consists of sixteen 32-bit words yi and can be arranged as a 4-by-4 matrix: u0 u1 u2 u3 u4 u5 u6 u7 u8 u9 u10 u11 u12 u13 u14 u15 To start the encryption process, the initial state of Salsa20 is constructed as follows. Eight 32-bit words are formed by the key k = [k0k1k2k3k4k5k6k7], two words indicate the stream position p = [p0 p1], two words come from the nonce n = [n0n1] and four words are a constant c = [c0c1c2c3]:
2.4 Practical Stream Ciphers 57 c0 k0 k1 k2 k3 c1 n0 n1 p0 p1 c2 k4 k5 k6 k7 c3 p can be seen as a counter indicating the position of the current 512-bit block within the range of all 264 512-bit blocks of the key stream. The constant c is given by the ASCII encoded string “expand 32-byte k”. If a 128-bit key consisting of 4 words k = [k0k1k2k3] is being used, the same key is concatenated to itself to form the required 8 words, i.e., k = [k0k1k2k3k0k1k2k3]. The core operation in Salsa20 is the quarter-round function QR(a, b, c, d) and is shown in Figure 2.8. It repeatedly applies three simple operations on 32-bit words: 32-bit addition modulo 232, 32-bit XOR and a constant 32-bit rotation by c positions to the left (ROT Lc). We note that the addition modulo 232 is simply a regular integer addition of two words, where the carry is ignored. The four-word output is computed from a four-word input by the quarter-round function QR as follows: b = b ⊕ ROT L7(a + d) c = c ⊕ ROT L9(b + a) d = d ⊕ ROT L13(c + b) a = a ⊕ ROT L18(d + c) Four quarter rounds form (not surprisingly) a round, and two consecutive rounds are called a double-round: In odd-numbered rounds, QR is applied to each of the four columns in the 4-by-4 matrix. In even-numbered rounds, QR is applied to each of the four rows. With an input (v0, v1, . . . , v15), the output (u0, u1, . . . , u15) of the first (odd) round of the double-round is computed as follows: (u0, u4, u8, u12) = QR(v0, v4, v8, v12) (u5, u9, u13, u1) = QR(v5, v9, v13, v1) (u10, u14, u2, u6) = QR(v10, v14, v2, v6) (u15, u3, u7, u11) = QR(v15, v3, v7, v11) The second round of the double-round yields the output (z0, z1, . . . , z15): (z0, z1, z2, z3) = QR(u0, u1, u2, u3) (z5, z6, z7, z4) = QR(u5, u6, u7, u4) (z10, z11, z8, z9) = QR(u10, u11, u8, u9) (z15, z12, z13, z14) = QR(u15, u12, u13, u14) Figure 2.9 shows a double round with the application of the quarter-round func- tion on the rows and colums. For encryption or decryption, 20 rounds or 10 double- rounds are applied.
58 2 Stream Ciphersa c db a c db <<< 7 <<< 9 <<<13 <<<18 Fig. 2.8 Quarter-round function QR(a, b, c, d) of Salsa20v0 v1 v2 v3 v12 v13 v14 v15 v4 v5 v6 v7 v8 v9 v10 v11 u0 u1 u2 u3 u12 u13 u14 u15 u4 u5 u6 u7 u8 u9 u10 u11 z0 z1 z2 z3 z12 z13 z14 z15 z4 z5 z6 z7 z8 z9 z10 z11 QR()QR() Fig. 2.9 Double-round function of Salsa20 Implementation Since Salsa20 is an ARX cipher, i.e., its internal structure uses only additions, ro- tation and XOR operations, it allows for very compact high-throughput software implementations. Salsa20 with 20 rounds requires approximately 4–14 cycles per byte on average for long streams, depending on the CPU type.
2.4 Practical Stream Ciphers 59 2.4.2 ChaCha ChaCha is another fast, software-oriented stream cipher, which was also developed by Bernstein in 2008. It follows the same basic design principles as Salsa20. The ci- pher can be configured with eight (ChaCha8), twelve (ChaCha12), or twenty rounds (ChaCha20). This section focuses on ChaCha20 with twenty rounds and a 256-bit key. Similarly to Salsa20, a 128-bit version exists too. Encryption and Decryption with ChaCha20 Encryption and decryption with ChaCha is performed in the same way as with Salsa20: ChaCha generates a 512-bit hash from its input values consisting of a key, a nonce and a block number. The hash value is XORed with a 512-bit plaintext or ciphertext for encryption or decryption, respectively. As with Salsa20, each 512- bit key stream block can be computed independently and encryption or decryption blocks can be computed in parallel. Let k be a 32-byte (256 bits) or 16-byte (128 bits) sequence. Let n be an 8- byte nonce. Let x be an l-byte message for some l ∈ 0, 1, . . . , 270. The ChaCha20 encryption of message x with a nonce n and a key k yields an l-byte ciphertext y: y = ChaCha20k(n) ⊕ x For decryption, the same key stream is used and XORed to the ciphertext, i.e., x = ChaCha20k(n) ⊕ y Core Function of ChaCha20 Similarly to Salsa20, ChaCha20’s initial state includes a 128-bit constant c, a 256- bit key k, a 64-bit counter p and a 64-bit nonce n, which we can arrange in a 4-by-4 matrix of 32-bit words: c0 c1 c2 c3 k0 k1 k2 k3 k4 k5 k6 k7 p0 p1 n0 n1 We note that there is a different order of the input values compared to Salsa20. The constant c is given by the ASCII encoded string “expand 32-byte k”. Like Salsa20, ChaCha20 uses a quarter-round function QR(a, b, c, d) on its 32-bit input values a, b, c and d. In the case of ChaCha20, it applies 4 additions modulo 232 and 4 XORs and 4 rotations repeatedly on the 32-bit state words. Compared to Salsa20, ChaCha applies the operations in a different order. Moreover, each word is
60 2 Stream Ciphers updated twice per round: a = a + b d = ROT L16(d ⊕ a) c = c + d b = ROT L12(b ⊕ c) a = a + b d = ROT L8(d ⊕ a) c = c + d b = ROT L7(b ⊕ c) Figure 2.10 shows the quarter-round function of ChaCha20.<<<12 <<<16 <<<8 <<<7 a c db a c db Fig. 2.10 Quarter-round function QR(a, b, c, d) of ChaCha20 Similarly to Salsa20, ChaCha20 performs 10 iterations of a double round. Again, each double round consists of two consecutive rounds, which in turn are each com-
2.4 Practical Stream Ciphers 61 posed of four quarter rounds QR. If we denote the input by (v0, v1, . . . , v15), the output (u0, u1, . . . , u15) of the first (odd) round of the double-round is computed on the columns of the input as follows: (u0, u4, u8, u12) = QR(v0, v4, v8, v12), (u1, u5, u9, u13) = QR(v1, v5, v9, v13), (u2, u6, u10, u14) = QR(v2, v6, v10, v14), (u3, u7, u11, u15) = QR(v3, v7, v11, v15). The second round of the double-round yields the intermediate output (z0, z1, . . . , z15): (z0, z5, z10, z15) = QR(v0, v5, v10, v15), (z1, z6, z11, z12) = QR(v1, v6, v11, v12), (z2, z7, z8, z13) = QR(v2, v7, v8, v13), (z3, z4, z9, z14) = QR(v3, v4, v9, v14). Implementation Compared to Salsa20, the rounds in ChaCha have an improved diffusion (cf. Sec- tion 3.1.1 for a discussion of diffusion in symmetric ciphers) while maintaining a similar performance. Compared to AES software implementations, ChaCha20 is about three times as fast on CPUs without specific AES accelerators. Like a Salsa20 round, a ChaCha round has 16 XORs, 16 additions, and 16 rotations of 32-bit words. ChaCha20 with 20 rounds requires approximately 4–15 cycles per byte on average for long streams, depending on the processor type. 2.4.3 Trivium Like Salsa20, Trivium also grew out of the eStream project. It was designed by Christophe De Canni`ere and Bart Preneel. In contrast to Salsa20, it is a hardware- oriented cipher. This means that hardware implementations of the cipher use few resources (i.e., logic gates) and should allow high encryption rates. Another differ- ence from Salsa20 is that it uses an 80-bit key. Trivium is based on a combination of three shift registers. Even though these are linear feedback shift registers, there are nonlinear components used to combine the registers, which prevents the attack against LFSRs that we studied in Section 2.3.2.
62 2 Stream Cipherssi key stream A B C 661 1 1 69 91 92 93 69 78 82 83 84 66 87 109 110 111 Fig. 2.11 Internal structure of the stream cipher Trivium Core Function of Trivium As shown in Figure 2.11, at the heart of Trivium are the three shift registers, A, B and C. The bit lengths of the registers are 93, 84 and 111, respectively. The XOR- sum of all three register outputs forms the key stream si. A specific feature of the cipher is that the output of each register is connected to the input of another register. Thus, the registers are arranged in a circle-like fashion. The cipher can be viewed as consisting of one circular register with a total length of 93 + 84 + 111 = 288. Each of the three registers has similar structure, as described below. The input of each register is computed as the XOR-sum of two bits: 1. For instance, the output of register A is part of the input of register B, as can be seen in Figure 2.11. 2. One register bit at a specific location is fed back to the input. The positions are given in Table 2.3. For instance, bit 69 of register A is fed back to its input. The output of each register is computed as the XOR-sum of three bits: The rightmost register bit. One register bit at a specific location is fed forward to the output. The positions are given in Table 2.3. For instance, bit 66 of register A is fed to its output. The output of a logical AND function whose inputs are two specific register bits. Again, the positions of the AND gate inputs are given in Table 2.3.
2.4 Practical Stream Ciphers 63 Table 2.3 Specification of Trivium register length feedback bit feedforward bit AND inputs A 93 69 66 91, 92 B 84 78 69 82, 83 C 111 87 66 109, 110 Alternatively, Trivium can be described using the following recursive formulae: ai ≡ ci−66 + ci−111 + ci−110 · ci−109 + ai−69 mod 2 bi ≡ ai−66 + ai−93 + ai−92 · ai−91 + bi−78 mod 2 ci ≡ bi−69 + bi−84 + bi−83 · bi−82 + ci−87 mod 2 The leftmost bits ai, bi and ci are the new inputs for the three registers. The key stream is then computed as the XOR sum of six register bits, cf. also Figure 2.11: si ≡ ai−66 + ai−93 + bi−69 + bi−84 + ci−66 + ci−111 mod 2 Note that the AND operation is equal to multiplication in modulo 2 arithmetic. This is in contrast to simple LFSRs, which only use XOR, i.e., modulo 2 addition. In fact, the feed forward paths involving the AND operations are crucial for the security of Trivium as they prevent attacks that exploit the linearity of the cipher, such as the one shown against plain LFSRs in Section 2.3.2. Encryption and Decryption with Trivium Almost all modern stream ciphers have two input parameters: a key k and an ini- tialization vector IV. The former is the regular key that is used in every symmetric cryptographic system. The IV serves as a randomizer and should take a new value for every encryption session, i.e., it should be a nonce (cf. the discussion of nonces in stream ciphers in Section 2.4.1). We look now at the details of running Trivium. Initialization Initially, an 80-bit key is loaded in the 80 leftmost locations of reg- ister A, and an 80-bit IV is loaded into the 80 leftmost locations of register B. All other register bits are set to zero with the exception of the three rightmost bits of register C, i.e., bits c109, c110 and c111, which are set to 1. Warm-Up Phase In the first phase, the cipher is clocked 4 × 288 = 1152 times, where 288 is the total length of all registers. No cipher output is generated. The warm-up phase is required to prevent an attacker from computing the key from the key stream. In other words, the warm-up phase is needed to randomize the cipher sufficiently. It makes sure that the key stream depends on both the key k and the IV in a way that cannot be predicted by an adversary.
64 2 Stream Ciphers Encryption (Decryption) Phase The bits produced hereafter, i.e., starting with the output bit of cycle 1153, form the key stream. As in every stream cipher, the key stream is XORed to the plainext (encryption) or ciphertext (decryption). Implementation An attractive feature of Trivium is its compactness, especially if implemented in hardware. It mainly consists of a 288-bit shift register and a few Boolean gates. It is estimated that a hardware implementation of the cipher occupies an area of between about 3500 and 5500 gate equivalences, depending on the degree of parallelization. (A gate equivalence is the chip area occupied by a 2-input NAND gate.) For instance, an implementation with 4000 gates computes the key stream at a rate of 16 bits/clock cycle. This is considerably smaller than most block ciphers such as AES and additionally is very fast. If we assume that this hardware design is clocked at 500 MHz, the encryption rate would be 16 bits × 500 MHz = 8 Gbit/s. In software, it is estimated that computing 8 output bits takes 12 cycles on a 1.5 GHz Intel CPU, resulting in a theoretical encryption rate of 1 Gbit/s. Security of Trivium At the time of writing no attack is known that is better than brute-force, i.e., that requires less then 280 steps. There are some attacks against weakened versions of Trivium. For instance, a key can be computed in 268 steps in case of a reduced ini- tialization phase of 799 iterations (rather than the 1152 steps in the Trivium specifi- cation). It should be kept in mind that Trivium was developed to be a very small and efficient cipher and is not intended for high-security applications. It can be spec- ulated that large nation-state attackers will be able to launch a brute-force attack against ciphers with 80 key bits in the not-too-distant future. 2.5 Discussion and Further Reading True Random Number Generation In this chapter we introduced different classes of RNGs, and showed that cryptographically secure pseudorandom number genera- tors are of central importance for stream ciphers. For other cryptographic applica- tions, true random number generators are often needed. For instance, TRNGs are used for the generation of cryptographic keys, which are then to be distributed among participants. Many stream ciphers and modes of operation rely on initial values that are often generated from TRNGs. Also, many protocols require nonces (numbers used only once), which may stem from a TRNG. All TRNGs need to ex- ploit some entropy source, i.e., some process which behaves truly randomly. Many TRNG designs have been proposed over the years. They can coarsely be classified as approaches that use specially designed hardware as a physical entropy source or as TRNGs that exploit existing components of computer systems as sources of randomness. Examples of the former are electronic circuits with random behavior,
2.5 Discussion and Further Reading 65 e.g., that are based on jitter in electronic circuits or several uncorrelated oscillators. Many modern CPUs are equipped with such hardware-based TRNGs. Reference [165, Chapter 5] contains a good survey on the topic. Examples of the latter are computer systems that measure the times between key strokes, system interrupts, or the arrival times of packets at network interfaces. Other possible entropy sources are checksum over memory or hard disc content. On Linux-like computer systems, the file /dev/random provides random bits that were collected in such a fashion. In all these cases, one has to be extremely careful to make sure that the noise source in fact has enough entropy. There are many examples of TRNG designs that turned out to have poor random behavior and which constitute a serious security weakness, depending on how they are used. There are tools available that test the statistical properties of TRNG output sequences [91, 195]. There are also standards with which TRNGs can be formally evaluated [121]. Hirstorical Remark on Stream Ciphers and the OTP In the literature, stream ciphers are often attributed to Gilbert Vernam who developed the concept in 1917, even though they were not called stream ciphers back at that time. He built an elec- tromechanical machine that automatically encrypted teletypewriter communication. The plaintext was fed into the machine as one paper tape, and the key stream as a second tape. This was the first time that encryption and transmission was auto- mated in one machine. Occasionally, one-time pads are also called Vernam ciphers. Vernam studied electrical engineering at Worcester Polytechnic Institute (WPI) in Massachusetts where, by coincidence, one of the authors of this book was a pro- fessor in the 1990s. Later on, Joseph Mauborgne discovered that the Vernam cipher is unbreakable if the key stream is truly random and not reused, i.e., it becomes an OTP. For further reading on Vernam’s machine, the book by Kahn [156] is recom- mended. More recently, it was discovered that the one-time pad had been invented 35 years prior to Vernam’s machine, in 1882 by a Sacramento banker named Frank Miller [34]. eSTREAM project As mentioned in the beginning of Section 2.4, the eSTREAM project [108] was initiated in 2004 to trigger the development of new stream ci- phers that are more secure and more efficient than many of the earlier stream cipher constructions. eSTREAM was organized by the European Network of Excellence in Cryptography (ECRYPT). In 2004, eSTREAM issued a call for new stream ci- phers and the selection process ended in 2008. The ciphers were divided into two “profiles”, depending on the intended application: Profile 1: Stream ciphers for software applications with high throughput require- ments. Profile 2: Stream ciphers for hardware applications with restricted resources such as limited storage, gate count or power consumption. Some cryptographers had emphasized the importance of including an authentication method, and hence two further profiles were also included to deal with ciphers that also provide authentication.
66 2 Stream Ciphers A total of 34 stream cipher candidates were submitted to the eSTREAM project. At the end of the project four software-oriented (Profile 1) ciphers were found to have desirable properties: HC-128, Rabbit, Salsa20/12 and SOSEMANUK. With re- spect to hardware-oriented ciphers (Profile 2), the following three ciphers were se- lected: Grain v1, MICKEY v2 and Trivium. The algorithm description, source code and the results of the four-year evaluation process are available online [108], and the official book provides more detailed information [219]. The official reference document for Salsa20 is [38] and [37] for ChaCha. It is important to keep in mind that ECRYPT is not a standardization body, so the status of the eSTREAM finalist ciphers cannot be compared to that of AES, which was initially standardized in the USA by NIST (cf. Section 4.1). Neverthe- less, the ChaCha ciphers, a variant of the Salsa algorithm, have been specified as an internet standard in RFC 7539 [201] and are part of the TLS cipher suite and of OpenSSH. Trivium has been standardized as a “lightweight cipher” in ISO/IEC 29192-3:2012 [151]. Other Stream Ciphers Even though many stream ciphers have been proposed over the years, many of the pre-eSTREAM algorithms are not as well scrutinized as the eSTREAM ciphers. The security of many of those older stream ciphers is unknown, and many of them have been broken. In the case of older software-oriented stream ciphers, arguably the best-investigated one is RC4 [217]. In 2015, the use of RC4 within Transport Layer Security (TLS) was prohibited due to severe security is- sues [6]. In the case of hardware-oriented ciphers, there is a wealth of LFSR-based al- gorithms. Again, many older proposed ciphers have been broken; see references [18, 126] for an introduction. Among the best-studied ones are the A5/1 and A5/2 algorithms, which are used in GSM mobile networks for voice encryption between cell phones and base stations. A5/1, which was the cipher used in most industrialized nations, had originally been kept secret but was reverse-engineered and published on the internet in 1998. The cipher was borderline secure at release time [49], whereas the weaker A5/2 has much more serious flaws [25]. Neither of the two ciphers is recommended based on today’s understanding of cryptanalysis. For 3G (or UMTS) mobile communication, a different cipher A5/3 (also named KASUMI) is used, but it is a block cipher.
2.6 Lessons Learned 67 2.6 Lessons Learned Stream ciphers are an important part of modern cryptography but are somewhat less widely used than block ciphers. The one-time pad is a provably secure symmetric cipher. However, it is highly impractical for most applications because the key length has to equal the message length. Stream ciphers sometimes require fewer resources, e.g., code size or chip area, for implementation than block ciphers, and they can be very fast. Secure and fast stream ciphers such as ChaCha20 can be built from functions that consist of the add-rotate-XOR operations. The requirements for a cryptographically secure pseudorandom number gener- ator are far more demanding than the requirements for pseudorandom number generators used in other fields of engineering such as testing or simulation. Single LFSRs make poor stream ciphers despite their good statistical properties. However, careful combination of several LFSRs can yield strong ciphers.
68 2 Stream Ciphers Problems 2.1. The stream cipher described in Definition 2.1.1 can easily be generalized to work in alphabets other than the binary one. For manual encryption, an especially useful one is a stream cipher that operates on letters. 1. Develop a scheme which operates with the letters A, B,. . ., Z, represented by the numbers 0,1,. . .,25. What does the key (stream) look like? What are the encryp- tion and decryption functions? 2. Decrypt the following ciphertext: bsaspp kkuosr which was encrypted using the key: rsidpy dkawoa 3. How was the young man murdered? 2.2. Assume we store a one-time key on a DVD with a capacity of 1 Gbyte. Discuss the real-life implications of a one-time pad (OTP) system. Address issues such as the life cycle of the key, storage of the key during the life cycle/after the life cycle, key distribution, generation of the key, etc. 2.3. Assume an OTP-like encryption with a short key of 128 bits. This key is then used periodically to encrypt large volumes of data. Describe how an attack works that breaks this scheme. 2.4. At first glance it seems as though an exhaustive key search is possible against an OTP system. Given is a short message, let’s say 5 ASCII characters represented by 40 bits, which was encrypted using a 40-bit OTP. Explain exactly why an exhaus- tive key search will not succeed even though sufficient computational resources are available. This is a paradox since we know that the OTP is unconditionally secure. That is, explain why a brute-force attack does not work. Note: You have to resolve the paradox! That means answers such as “The OTP is unconditionally secure and therefore a brute-force attack does not work” are not valid. 2.5. The OTP can be used to encrypt data of arbitrary length by encrypting binary symbols xi ∈ {0, 1}. Decrypt the following ciphertext by hand. The ciphertext is given in hexadecimal notation: 26 34 05 18 0c 06 07 15 1c 2a 13 3c 0c 23 04 27 07 27 18 The key is given by: 6a 51 71 6b 49 68 64 67 65 5a 67 68 64 4a 77 65 68 48 73 Compute the plaintext, which is encoded in ASCII symbols. 2.6. The OTP offers provable security. Describe two major drawbacks of the OTP that render it impractical for most applications such as encryption of emails or in- stant messenging.
2.6 Problems 69 2.7. We will now analyze a pseudorandom number sequence generated by an LFSR of degree 3 characterized by (p2 = 1, p1 = 0, p0 = 1). 1. What is the sequence generated from the initialization vector: (s2 = 1, s1 = 0, s0 = 0)? 2. What is the sequence generated from the initialization vector: (s2 = 0, s1 = 1, s0 = 1)? 3. How are the two sequences related? 2.8. Assume we have a stream cipher whose period is quite short. We happen to know that the period is 150–200 bits in length. We assume that we do not know anything else about the internals of the stream cipher. In particular, we should not assume that it is a simple LFSR. For simplicity, assume that English text in ASCII format is being encrypted. Describe in detail how such a cipher can be attacked. Specify exactly what Oscar has to know in terms of plaintext/ciphertext, and how he can decrypt all ciphertext. 2.9. Compute the first two output bytes of the LFSR of degree 8 and the feedback polynomial from Table 2.2 where the initialization vector has the value FF in hex- adecimal notation. 2.10. In this problem we will study LFSRs in somewhat more detail. LFSRs come in three flavors: LFSRs which generate a maximum-length sequence. These LFSRs are based on primitive polynomials. LFSRs which do not generate a maximum-length sequence but whose sequence length is independent of the initial value of the register. These LFSRs are based on irreducible polynomials that are not primitive. (Note that all primitive poly- nomials are also irreducible.) LFSRs which do not generate a maximum-length sequence and whose sequence length depends on the initial values of the register. These LFSRs are based on reducible polynomials. We will study examples in the following. Determine all sequences generated by the following three polynomials: 1. x4 + x + 1 2. x4 + x2 + 1 3. x4 + x3 + x2 + x + 1 Draw the corresponding LFSR for each of the three polynomials. Which of the polynomials is primitive, which is only irreducible, and which one is reducible? Note that the lengths of all sequences generated by each of the LFSRs should always add up to 2m − 1.
70 2 Stream Ciphers 2.11. Given is a stream cipher based on a single LFSR as key stream generator. The LFSR has a degree of 256. 1. How many plaintext/ciphertext bit pairs are needed to launch a successful attack? 2. Describe all steps of the attack in detail and develop the formulae that need to be solved. 3. What is the key in this system? Why doesn’t it make sense to use the initial contents of the LFSR as the key or as part of the key? 2.12. We conduct a known-plaintext attack on an LFSR-based stream cipher. We know that the plaintext sent was: 1001 0010 0110 1101 1001 0010 0110 By tapping the channel we observe the following stream: 1011 1100 0011 0001 0010 1011 0001 1. What is the degree m of the key stream generator? 2. What is the initialization vector? 3. Determine the feedback coefficients of the LFSR. 4. Draw a circuit diagram and verify the output sequence of the LFSR. 2.13. We want to perform an attack on another LFSR-based stream cipher. In order to process letters, each of the 26 uppercase letters and the numbers 0, 1, 2, 3, 4, 5 are represented by a 5-bit vector according to the following mapping: A ↔ 0 = 000002 ... Z ↔ 25 = 110012 0 ↔ 26 = 110102 ... 5 ↔ 31 = 111112 We happen to know the following facts about the system: The degree of the LFSR is m = 6. Every message starts with the header WPI. We observe now on the channel the following message (the fourth symbol is a zero): j5a0edj2b 1. What is the initialization vector? 2. What are the feedback coefficients of the LFSR? 3. Write a program in your favorite programming language which generates the whole sequence, and find the whole plaintext. 4. Where does the thing after WPI live? 5. What type of attack did we perform?
2.6 Problems 71 2.14. In this problem we will look at pseudorandom number generators based on a linear congruential generator (LCG). As we have seen in Section 2.2.1, an LCG is given by the equations: z0 ≡ seed zi+1 ≡ a · zi + b mod m , i = 0, 1, ... We assume that the modulus m is public and that the key is formed by the parameters seed, a and b. We consider now the problem that arises if we use an LCG as key stream genera- tor. We assume the stream cipher is used for encrypting images given in GIF format. The key stream zi encrypts a plaintext xi as follows: yi ≡ xi + zi mod m. Since GIF files consist of 8-bit values, we need at least 256 possible values for xi. Thus, the prime modulus m = 257 is a good choice. Now, assume that the first six bytes in the header of a GIF-image file consists of the letters GIF89a, where each letter is encoded as an 8-bit ASCII character. An attacker who obtains an encrypted GIF file finds the following ciphertexts at the beginning of the file: y1 = 32, y2 = 166 and y3 = 87, which correspond to the plaintext bytes x1 = G, x2 = I and x3 = F. 1. Describe how an attacker can compute the parameters a, b as well as the seed with these three plaintext bytes. Compute the parameters a, b and the seed. 2. What is this attack called? What are the prerequisites for a successful attack? 2.15. The linear congruential generator described in Section 2.2.1 can be extended such that each new element zi+1 of the key stream is computed from the two pre- vious elements zi and zi−1. In this case, two seed values z0 and z1 as well as three parameters a, b and c along with the modulus m are needed. The equation of the generator is given as: zi+1 ≡ a · zi + b · zi−1 + c mod m The key stream zi is used to encrypt letters given in 8-bit ASCII code on a character-by-character basis. That means for each key stream value zi, one plain- text character xi is encrypted as yi ≡ xi + zi mod m. Oscar, the attacker, eavesdrops on the communication and happens to know that the ciphertext contains the name ALICE, starting at position i. Oscar observes now the following ciphertext symbols on the channel: yi = 69, yi+1 = 47, yi+2 = 3, yi+3 = 88, yi+4 = 217 He also knows that the modulus m = 257 is being used. Show how Oscar can com- pute the parameters a, b and c.
72 2 Stream Ciphers 2.16. We want to use the Salsa20 algorithm with a 256-bit key for encryption. Let the key k and the nonce n consist of only zeros and let the stream position start at zero. Provide the 512-bit initial state of Salsa20. 2.17. Let us now have a closer look at the quarter-round function of ChaCha, i.e., QR(a, b, c, d). What is the output of QR for the following input? a = 0x00000001 b = 0x00000000 c = 0x00000000 d = 0x00000000 2.18. Assume the IV and the key of Trivium each consist of 80 all-zero bits. Com- pute the first 70 bits s1, . . . , s70 during the warm-up phase of Trivium. Note that these are only internal bits, which are not used for encryption since the warm-up phase lasts for 1152 clock cycles.
Chapter 3 The Data Encryption Standard (DES) and Alternatives The Data Encryption Standard, or DES, was conceived in the early 1970s and can arguably be considered the first modern encryption algorithm. It was the most pop- ular block cipher in the 1980s and 1990s. Even though DES in its basic form is nowadays not secure because of the small key space, its variant 3DES or triple DES is still in use in the 2020s (cf. Section 3.7.2). 3DES simply encrypts data three times in a row with DES. The design principles of DES have inspired many current ciphers and, hence, studying it helps us to understand many other symmetric algorithms. In this chapter you will learn: The design process of DES, which is very helpful for understanding the technical and political evolution of modern cryptography Basic design ideas of block ciphers, including confusion and diffusion, which are important properties of all modern block ciphers The internal structure of DES, including Feistel networks, S-boxes and the key schedule Security analysis of DES Alternatives to DES, including 3DES and the lightweight block cipher PRESENT 73 C. Paar et al., Understanding Cryptography, https://doi.org/10.1007/978-3-662-69007-9_3 © The Editor(s) (if applicable) and The Author(s), under exclusive license to Springer-Verlag GmbH, DE, part of Springer Nature 2024
74 3 The Data Encryption Standard (DES) and Alternatives 3.1 Introduction to DES In 1972 a mildly revolutionary act was performed by the U.S. National Bureau of Standards (NBS), which is now called the National Institute of Standards and Tech- nology (NIST): the NBS initiated a request for proposals for a standardized cipher in the USA. The idea was to find a single secure cryptographic algorithm which could be used for a variety of applications. Up to this point in time governments had always considered cryptography, and in particular cryptanalysis, so crucial for na- tional security that it had to be kept secret. However, by the early 1970s the demand for encryption for commercial applications such as banking had become so pressing that it could not be ignored without economic consequences. The NBS received the most promising candidate in 1974 from a team of cryp- tographers working at IBM. The algorithm IBM submitted was based on the cipher Lucifer. Lucifer was a family of ciphers developed by Horst Feistel in the late 1960s, and was one of the first instances of block ciphers operating on digital data. Lucifer is a Feistel cipher which encrypts blocks of 64 bits using a key size of 128 bits. In order to investigate the security of the submitted ciphers, the NBS requested the help of the National Security Agency (NSA), which did not even admit its existence at that point in time1. It seems certain that the NSA influenced changes to the ci- pher, which was rechristened DES. One of the changes that occurred was that DES is specifically designed to withstand differential cryptanalysis, an attack not known to the public until 1990. It is not clear whether the IBM team developed the knowl- edge about differential cryptanalysis by themselves or whether they were guided by the NSA. Allegedly, the NSA also convinced IBM to reduce the Lucifer key length of 128 bits to 56 bits, which made the cipher much more vulnerable to brute-force attacks. The NSA involvement worried some people because it was feared that a secret backdoor, i.e., a mathematical property with which DES could be broken but which is only known to the NSA, might have been the real reason for the modifications. An- other major complaint was the reduction of the key size. Some people conjectured that the NSA would be able to search through a key space of 256, thus breaking it by brute-force. In later decades, most of these concerns turned out to be unfounded. Section 3.5 provides more information about real and perceived security weaknesses of DES. Despite all the criticism and concerns, in 1977 the NBS finally released all spec- ifications of the modified IBM cipher to the public as Data Encryption Standard (FIPS PUB 46). Even though the cipher is described down to the bit level in the standard, the motivation for parts of the DES design (the so-called design criteria), especially the choice of the substitution boxes, was never officially released. With the rapid increase in personal computers in the early 1980s and all specifica- tions of DES being publicly available, it became easier to analyze the inner structure of the cipher. During this period, the academic cryptography research community also grew and DES underwent major scrutiny. However, no serious weaknesses were 1 A standard joke in the cryptographic community was that NSA stands for “no such agency”.
3.1 Introduction to DES 75 found until 1990. Originally, DES was only standardized for 10 years, until 1987. Due to the wide use of DES and the lack of security weaknesses, NIST reaffirmed the federal use of the cipher until 1999, when it was finally replaced by the Advanced Encryption Standard (AES). 3.1.1 Confusion and Diffusion Before we consider the details of DES, it is instructive to look at basic operations that can be applied in order to achieve strong encryption. According to the famous information theorist Claude Shannon, there are two primitive operations with which strong encryption algorithms can be built [232]: 1. Confusion is an encryption operation where the relationship between key and ciphertext is obscured. Today, a common element for achieving confusion is sub- stitution, which is found in both DES and AES. 2. Diffusion is an encryption operation where the influence of one plaintext sym- bol is spread over many ciphertext symbols with the goal of hiding statistical properties of the plaintext. A simple diffusion element is the bit permutation, which is used frequently within DES. AES uses the more advanced MixColumn operation. Ciphers which only perform confusion, such as the Shift Cipher (cf. Section 1.4.3) or the World War II encryption machine Enigma, are not secure. Neither are ciphers which only perform diffusion. However, through the concatenation of such oper- ations, a strong cipher can be built. The idea of concatenating several encryption operations was also proposed by Shannon. Such ciphers are known as product ci- phers. All of today’s block ciphers are product ciphers as they consist of rounds which are applied repeatedly to the data (Figure 3.1). Modern block ciphers possess excellent diffusion properties. On a cipher level this means that changing one bit of plaintext results on average in changing half the output bits, i.e., the second ciphertext looks statistically independent of the first one. This is an important property to keep in mind when dealing with block ciphers. We demonstrate this behavior with the following simple example. Example 3.1. Let’s assume a toy block cipher with a block length of 8 bits. Encryp- tion of two plaintexts x1 and x2, which differ only by one bit, should roughly result in the situation shown in Figure 3.2. Note that modern block ciphers have block lengths of 64 or 128 bits but they show exactly the same behavior if one input bit is flipped.
76 3 The Data Encryption Standard (DES) and Alternatives Fig. 3.1 Principle of an N round product cipher, where each round performs a con- fusion and a diffusion operationBlock Cipher = 0110 11002 y = 1011 10011 x = 0000 10112 x = 0010 10111 y Fig. 3.2 Principle of diffusion of a block cipher: A one-bit change in the input leads to statistically independent outputs 3.2 Overview of the DES Algorithm DES is a cipher that encrypts blocks of length 64 bits with a key of size of 56 bits (Figure 3.3). DES is a symmetric cipher, i.e., the same key is used for encryption and decryp- tion. DES is, like virtually all modern block ciphers, a round-based algorithm. For each block of plaintext, encryption is handled in 16 rounds, which all perform the identical operation. Figure 3.4 shows the round structure of DES. In every round a different subkey is used and all subkeys ki are derived from the main key k. Let’s now have a more detailed look at the internals of DES, as shown in Fig- ure 3.5. The structure in the figure is called a Feistel network. It can lead to very strong ciphers if carefully designed. Feistel networks are used in other, but certainly
3.2 Overview of the DES Algorithm 7764 64 k 56 x y DES Fig. 3.3 DES block cipherPermutation Initial Permutation Final Round 16 Encryption Round 1 Encryption y x k1 k16 k Fig. 3.4 Round structure of DES not in all, modern block ciphers too. (In fact, AES is not a Feistel cipher.) In ad- dition to its potential cryptographic strength, one advantage of Feistel networks is that encryption and decryption are almost the same operation. Decryption requires only a reversed key schedule, which is an advantage in software and hardware im- plementations. We discuss the Feistel network in the following.
78 3 The Data Encryption Standard (DES) and Alternatives64 k xDES ( )y = Round 1 56 56 16k 48 1k 48 1R1L Plaintext x IP(x) Initial Permutation 16R16L 15R15L 0R0L . . . . . . . . . . . . Ciphertext IP ( ) Transform 16 Transform 1 PC−1 Key k Round 16 −1 Final Permutation 32 32 3232 f 32 32 3232 f 64 Fig. 3.5 The Feistel structure of DES
3.3 Internal Structure of DES 79 After the initial bitwise permutation IP of a 64-bit plaintext block x, the plaintext is split into two halves L0 and R0. These two 32-bit halves are the input to the Feistel network, which consists of 16 rounds. The right half Ri is fed into the function f . The output of the f function is XORed (as usual, denoted by the symbol ⊕) with the 32-bit left half Li. Finally, the right and left halves are swapped. This process repeats in the next round and can be expressed as: Li = Ri−1 Ri = Li−1 ⊕ f (Ri−1, ki) where i = 1, . . . , 16. After round 16, the 32-bit halves L16 and R16 are swapped again, and the final permutation IP−1 is the last operation of DES. As the notation suggests, the final permutation IP−1 is the inverse of the initial permutation IP. In each round, a round key ki is derived from the main 56-bit key using what is called the key schedule. It is crucial to note that the Feistel structure only encrypts (decrypts) half of the input bits each round, namely the left half of the input. The right half is copied to the next round unchanged. In particular, the right half is not encrypted with the f function. In order to get a better understanding of the working of Feistel ciphers, the following interpretation is helpful: Think of the f function as a pseudorandom gen- erator with the two input parameters Ri−1 and ki. The output of the pseudorandom generator is then used to encrypt the left half Li−1 with an XOR operation. As we saw in Chapter 2, if the output of the f function is not predictable for an attacker, this results in a strong encryption method. The two basic properties of ciphers mentioned in Section 3.1.1, i.e., confusion and diffusion, are realized within the f function. In order to thwart advanced analyt- ical attacks, the f function must be designed extremely carefully. Once the f func- tion has been designed securely, the security of a Feistel cipher increases with the number of key bits used and the number of rounds. Before we discuss all components of DES in detail, here is an algebraic descrip- tion of the Feistel network for the mathematically inclined reader. The Feistel struc- ture of each round bijectively maps a block of 64 input bits to 64 output bits (i.e., every possible input is mapped uniquely to exactly one output, and vice versa). This mapping remains bijective for some arbitrary function f , i.e., even if the embedded function f is not bijective itself. In the case of DES, the function f is in fact a sur- jective many-to-one mapping. It uses nonlinear building blocks and maps 32 input bits to 32 output bits using a 48-bit round key ki, with 1 ≤ i ≤ 16. 3.3 Internal Structure of DES The structure of DES as depicted in Figure 3.5 shows the internal functions, which we will discuss in this section. The building blocks are the initial and final permu- tation, the actual DES rounds with their core, the f function, and the key schedule.
80 3 The Data Encryption Standard (DES) and Alternatives 3.3.1 Initial and Final Permutation As shown in Figures 3.6 and 3.7, the initial permutation IP and the final permuta- tion IP−1 are bitwise permutations. A bitwise permutation can be viewed as simple crosswiring. Interestingly, permutations can be very easily implemented in hard- ware but are not particularly fast in software. Note that both permutations do not increase the security of DES at all. The rationale for these two permutations in DES is purely implementational: The original purpose was to make it easier to arrange the plaintext and ciphertext bits in a bytewise manner to make data fetches easier for 8-bit data buses, which were the state-of-the-art register size in the early 1970s.. . . . . . . . . . . . . . . 1 1 50 IP(x) x 6458 402 Fig. 3.6 Examples of bit swaps in the initial permutation−1 64 . . . 5850 . . . . . . 1 21 40 . . . . . . IP (z) z Fig. 3.7 Examples for bit swaps in the final permutation
3.3 Internal Structure of DES 81 The details of the permutation IP are given in Table 3.1. This table, like all other tables in this chapter, should be read from left to right, top to bottom. The table indicates that input bit 58 is mapped to output position 1, input bit 50 is mapped to the second output position, and so forth. The final permutation IP−1 performs the inverse operation of IP as shown in Table 3.2. Table 3.1 Initial permutation IP IP 58 50 42 34 26 18 10 2 60 52 44 36 28 20 12 4 62 54 46 38 30 22 14 6 64 56 48 40 32 24 16 8 57 49 41 33 25 17 9 1 59 51 43 35 27 19 11 3 61 53 45 37 29 21 13 5 63 55 47 39 31 23 15 7 Table 3.2 Final permutation IP−1 IP−1 40 8 48 16 56 24 64 32 39 7 47 15 55 23 63 31 38 6 46 14 54 22 62 30 37 5 45 13 53 21 61 29 36 4 44 12 52 20 60 28 35 3 43 11 51 19 59 27 34 2 42 10 50 18 58 26 33 1 41 9 49 17 57 25 3.3.2 The f Function As mentioned earlier, the f function plays a crucial role for the security of DES. In round i it takes the right half Ri−1 of the output of the previous round and the current round key ki as input. The output of the f function is used as an XOR-mask for encrypting the left half input bits Li−1. The structure of the f function is shown in Figure 3.8. First, the 32-bit input is ex- panded to 48 bits by partitioning the input into eight 4-bit blocks and by expanding each block to 6 bits. This happens inside the E-box, which is a special type of per- mutation. The first block consists of the bits (1, 2, 3, 4), the second one of (5, 6, 7, 8), etc. The expansion to six bits can be seen in Figure 3.9. As shown in Table 3.3, exactly 16 of the 32 input bits appear twice in the output. However, an input bit never appears twice in the same 6-bit output block. The ex- pansion box increases the diffusion behavior of DES since some input bits influence two different output locations. Next, the 48-bit result of the expansion is XORed with the round key ki, and the eight 6-bit blocks are fed into eight different substition boxes, which are commonly referred to as S-boxes. Each S-box is a lookup table that maps a 6-bit input to a 4-bit output. Larger tables would have been cryptographically better but they also become much larger; eight 4-by-6 tables were probably close to the maximum size that could fit on a single integrated circuit in the early 1970s, when DES was designed. Each S-box contains 26 = 64 entries, which are typically represented by a table with 16 columns and 4 rows. Each entry is a 4-bit value. All S-boxes are listed in Tables 3.4 to 3.11. Note that all S-boxes are different. The tables are to be read as indicated
82 3 The Data Encryption Standard (DES) and AlternativesPermutation Expansion i−1E(R ) R i−1 S1 S2 S 3 S4 S5 S6 S7 S8 6 6 6 6 6 6 4 4 4 4 4 4 4 4 32 32 ki 32 48 48 48 6 6 P Fig. 3.8 Block diagram of the f function48 2 3 5 326 3 4 5 8 9 1110 47 1 4 7 8 9 142 . . . . . .1 6 7 12 13 Fig. 3.9 Examples of bit swaps in the expansion function E
3.3 Internal Structure of DES 83 Table 3.3 Expansion permutation E E 32 1 2 3 4 5 4 5 6 7 8 9 8 9 10 11 12 13 12 13 14 15 16 17 16 17 18 19 20 21 20 21 22 23 24 25 24 25 26 27 28 29 28 29 30 31 32 1 in Figure 3.10: the most significant bit (MSB) and the least significant bit (LSB) of each 6-bit input select the row of the table, while the four inner bits select the column. The integers 0,1,. . . ,15 of each entry in the table represent the decimal notation of a 4-bit value. Example 3.2. The S-box input b = (100101)2 indicates the row 112 = 3 (i.e., fourth row, numbering starts with 002) and the column 00102 = 2 (i.e., the third column). If the input b is fed into S-box 1, the output is S1(37 = 1001012) = 8 = 10002.fourth row 1 0 0 1 0 1 11 0 0 1 0 third column Fig. 3.10 Example of the decoding of the input 1001012 by S-box 1 Table 3.4 S-box S1 S1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 14 04 13 01 02 15 11 08 03 10 06 12 05 09 00 07 1 00 15 07 04 14 02 13 01 10 06 12 11 09 05 03 08 2 04 01 14 08 13 06 02 11 15 12 09 07 03 10 05 00 3 15 12 08 02 04 09 01 07 05 11 03 14 10 00 06 13
84 3 The Data Encryption Standard (DES) and Alternatives Table 3.5 S-box S2 S2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 15 01 08 14 06 11 03 04 09 07 02 13 12 00 05 10 1 03 13 04 07 15 02 08 14 12 00 01 10 06 09 11 05 2 00 14 07 11 10 04 13 01 05 08 12 06 09 03 02 15 3 13 08 10 01 03 15 04 02 11 06 07 12 00 05 14 09 Table 3.6 S-box S3 S3 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 10 00 09 14 06 03 15 05 01 13 12 07 11 04 02 08 1 13 07 00 09 03 04 06 10 02 08 05 14 12 11 15 01 2 13 06 04 09 08 15 03 00 11 01 02 12 05 10 14 07 3 01 10 13 00 06 09 08 07 04 15 14 03 11 05 02 12 Table 3.7 S-box S4 S4 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 07 13 14 03 00 06 09 10 01 02 08 05 11 12 04 15 1 13 08 11 05 06 15 00 03 04 07 02 12 01 10 14 09 2 10 06 09 00 12 11 07 13 15 01 03 14 05 02 08 04 3 03 15 00 06 10 01 13 08 09 04 05 11 12 07 02 14 Table 3.8 S-box S5 S5 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 02 12 04 01 07 10 11 06 08 05 03 15 13 00 14 09 1 14 11 02 12 04 07 13 01 05 00 15 10 03 09 08 06 2 04 02 01 11 10 13 07 08 15 09 12 05 06 03 00 14 3 11 08 12 07 01 14 02 13 06 15 00 09 10 04 05 03 Table 3.9 S-box S6 S6 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 12 01 10 15 09 02 06 08 00 13 03 04 14 07 05 11 1 10 15 04 02 07 12 09 05 06 01 13 14 00 11 03 08 2 09 14 15 05 02 08 12 03 07 00 04 10 01 13 11 06 3 04 03 02 12 09 05 15 10 11 14 01 07 06 00 08 13 Table 3.10 S-box S7 S7 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 04 11 02 14 15 00 08 13 03 12 09 07 05 10 06 01 1 13 00 11 07 04 09 01 10 14 03 05 12 02 15 08 06 2 01 04 11 13 12 03 07 14 10 15 06 08 00 05 09 02 3 06 11 13 08 01 04 10 07 09 05 00 15 14 02 03 12 The S-boxes are the core of DES in terms of cryptographic strength. They are the only nonlinear element in the algorithm and provide confusion.
3.3 Internal Structure of DES 85 Table 3.11 S-box S8 S8 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 13 02 08 04 06 15 11 01 10 09 03 14 05 00 12 07 1 01 15 13 08 10 03 07 04 12 05 06 11 00 14 09 02 2 07 11 04 01 09 12 14 02 00 06 10 13 15 03 05 08 3 02 01 14 07 04 10 08 13 15 12 09 00 03 05 06 11 Even though the entire specification of DES was released by NBS/NIST in 1977, the motivation for the choice of the S-box tables was never completely revealed. This often gave rise to speculation, in particular with respect to the possible exis- tence of a secret backdoor or some other intentionally constructed weakness which could be exploited by the NSA. However, today we know that the S-boxes were designed according to the criteria listed below. 1. Each S-box has six input bits and four output bits. 2. No single output bit should be too close to a linear combination of the input bits. 3. If the lowest and the highest bits of the input are fixed and the four middle bits are varied, each of the possible 4-bit output values must occur exactly once. 4. If two inputs to an S-box differ in exactly one bit, their outputs must differ in at least two bits. 5. If two inputs to an S-box differ in the two middle bits, their outputs must differ in at least two bits. 6. If two inputs to an S-box differ in their first two bits and are identical in their last two bits, the two outputs must be different. 7. For any nonzero 6-bit difference between inputs, no more than 8 of the 32 pairs of inputs exhibiting that difference may result in the same output difference. 8. A collision (zero output difference) at the 32-bit output of the eight S-boxes is only possible for three adjacent S-boxes. Note that some of these design criteria were not revealed until the 1990s. More information about the issue of the secrecy of the design criteria is found in Sec- tion 3.5. The S-boxes are the most crucial elements of DES because they introduce non- linearity into the cipher, i.e., S(a) ⊕ S(b) 6 = S(a ⊕ b) Without a nonlinear building block, an attacker could express the DES input and out- put with a system of linear equations where the key bits are the unknowns. Such sys- tems can easily be solved, a fact that was used in the LFSR attack in Section 2.3.2. However, the S-boxes were carefully designed to also thwart advanced mathematical attacks, in particular differential cryptanalysis. Interestingly, differential cryptanal- ysis was first discovered in the research community in 1990. At this point, the IBM team declared that the attack was known to the designers at least 16 years earlier, and that DES was especially designed to withstand differential cryptanalysis.
86 3 The Data Encryption Standard (DES) and Alternatives Finally, the 32-bit output of the S-Boxes is permuted bitwise according to the P permutation, which is given in Table 3.12. Unlike the initial permutation IP and its inverse IP−1, the permutation P has an important cryptographic purpose. It in- troduces diffusion because the four output bits of each S-box are permuted in such a way that they affect several different S-boxes in the following round. The diffu- sion caused by the expansion E and the permutation P together with the confusion caused by the S-boxes guarantee that each of the 64 bits at the end of the fifth round is a function of every plaintext bit and every key bit. This behavior is known as the avalanche effect. Table 3.12 The permutation P within the f function P 16 7 20 21 29 12 28 17 1 15 23 26 5 18 31 10 2 8 24 14 32 27 3 9 19 13 30 6 22 11 4 25 3.3.3 Key Schedule The key schedule derives 16 round keys ki, each consisting of 48 bits, from the original 56-bit key. Another term for round key is subkey. First, note that the DES input key is often stated to be 64 bits, where every eighth bit is used as an odd parity bit over the preceding seven bits, as shown in Figure 3.11. It is not quite clear why DES was specified that way. In any case, the eight parity bits are not actual key bits and do not increase the security. DES is a 56-bit cipher, not a 64-bit one! At the beginning of the key schedule, the 64-bit key is reduced to 56 bits by ignoring every eighth bit, i.e., the parity bits are stripped in the initial PC–1 permu- tation, cf. Figure 3.12. Again, the parity bits certainly do not increase the key space. The name PC–1 stands for “permuted choice one”. The exact bit connections that are realized by PC–1 are given in Table 3.13. Table 3.13 Initial key permutation PC–1 PC–1 57 49 41 33 25 17 9 1 58 50 42 34 26 18 10 2 59 51 43 35 27 19 11 3 60 52 44 36 63 55 47 39 31 23 15 7 62 54 46 38 30 22 14 6 61 53 45 37 29 21 13 5 28 20 12 4
3.3 Internal Structure of DES 87. . . LSB MSB 1 P = parity bit 64 P P 7 17 Fig. 3.11 Location of the eight parity bits for a 64-bit input key The resulting 56-bit key is split into two halves C0 and D0, and the actual key schedule starts as shown in Figure 3.12. The two 28-bit halves are cyclically shifted (i.e., rotated) left by one or two bit positions depending on the round i according to the following rule: In rounds i = 1, 2, 9, 16, the two halves are rotated left by one bit. In the other rounds where i 6 = 1, 2, 9, 16, the two halves are rotated left by two bits. Note that the rotations take place within the left and the right half. The total number of rotation positions is 4 · 1 + 12 · 2 = 28. This leads to the interesting property that C0 = C16 and D0 = D16. This is very useful for the decryption key schedule where the subkeys have to be generated in reversed order, as we will see in Section 3.4. To derive the 48-bit round keys ki, the two halves are permuted bitwise again with PC–2, which stands for “permuted choice 2”. PC–2 permutes the 56 input bits coming from Ci and Di and ignores 8 of them. The exact bit connections of PC–2 are given in Table 3.14. Table 3.14 Round key permutation PC–2 PC–2 14 17 11 24 1 5 3 28 15 6 21 10 23 19 12 4 26 8 16 7 27 20 13 2 41 52 31 37 47 55 30 40 51 45 33 48 44 49 39 56 34 53 46 42 50 36 29 32 Note that every round key is a selection of 48 permuted bits of the input key k. The key schedule is merely a method of realizing the 16 permutations systemati- cally. Especially in hardware, the key schedule is very easy to implement. The key schedule is also designed so that each of the 56 key bits is used in different round keys; each bit is used in approximately 14 of the 16 round keys.
88 3 The Data Encryption Standard (DES) and AlternativesTransform 1 Transform 16 16 1 . . . . . . . . . PC − 1 PC − 2 PC − 2 1 LS 00 1 11 22 1616 1616 DC LS DC LS LS LS LS DCk k k 64 56 56 56 28 48 48 2828 2828 28 Fig. 3.12 Key schedule for DES encryption 3.4 Decryption One advantage of DES is that decryption is essentially the same function as encryp- tion. This is a property of all Feistel ciphers. Figure 3.13 shows a block diagram for DES decryption. Compared to encryption, only the key schedule is reversed, i.e., in decryption round 1, subkey 16 is needed; in round 2, subkey 15; etc. Thus, when in decryption mode, the key schedule algorithm has to generate the round keys as the sequence k16, k15, . . . , k1.
3.4 Decryption 891k 16k k xDES ( )y = Round 1 56 56 1R1L IP(x) Initial Permutation 16R16L 15R 0R0L dd dd dd dd Message −1 k yDES ( )x = 48 48 Transform 1 Transform 16 Ciphertext IP ( ) PC−1 Key k Round 16 −1 Final Permutation 32 32 3232 f 32 32 3232 15 f L Fig. 3.13 DES decryption Reversed Key Schedule The first question that we have to clarify is how, given the initial DES key k, can we efficiently generate k16? Note that we saw above that C0 = C16 and D0 = D16.
90 3 The Data Encryption Standard (DES) and Alternatives Hence k16 can be directly derived from PC–1. k16 = PC–2(C16, D16) = PC–2(C0, D0) = PC–2(PC–1(k)) To compute k15 we need the intermediate variables C15 and D15, which can be de- rived from C16, D16 through cyclic right shifts (RS): k15 = PC–2(C15, D15) = PC–2(RS2(C16), RS2(D16)) = PC–2(RS2(C0), RS2(D0)) The subsequent round keys k14, k13, . . . , k1 are derived via right shifts in a similar fashion. The number of bits shifted right for each round key in decryption mode are: In decryption round 1, the key is not rotated. In decryption rounds 2, 9 and 16 the two halves are rotated right by one bit. In the other rounds 3, 4, 5, 6, 7, 8, 10, 11, 12, 13, 14 and 15 the two halves are rotated right by two bits. Figure 3.14 shows the reversed key schedule for decryption. Decryption in Feistel Networks We have not yet addressed the core question: Why is the decryption function es- sentially the same as the encryption function? The basic idea is that the decryption function reverses the DES encryption in a round-by-round manner. That means that decryption round 1 reverses encryption round 16, decryption round 2 reverses en- cryption round 15, and so on. Let’s first look at the initial stage of decryption by looking at Figure 3.13. Note that the right and left halves are swapped in the last round of DES: (Ld 0 , Rd 0 ) = IP(Y ) = IP(IP−1(R16, L16)) = (R16, L16) And thus: Ld 0 = R16 Rd 0 = L16 = R15 Note that all variables in the decryption routine are marked with the superscript d, whereas the encryption variables do not have superscripts. The derived equation simply says that the input of the first round of decryption is the output of the last round of encryption because the final and initial permutations cancel each other out.
3.4 Decryption 91Transform 16 Fig. 3.14 Reversed key schedule for decryption of DES We will now show that the first decryption round reverses the last encryption round. For this, we have to express the output values (Ld 1 , Rd 1 ) of the first decryption round 1 in terms of the input values of the last encryption round (L15, R15) . The first one is easy: Ld 1 = Rd 0 = L16 = R15 We now look at how Rd 1 is computed: Rd 1 = Ld 0 ⊕ f (Rd 0 , k16) = R16 ⊕ f (L16, k16) = [L15 ⊕ f (R15, k16)] ⊕ f (R15, k16) = L15 ⊕ [ f (R15, k16) ⊕ f (R15, k16)] = L15 The crucial step is shown in the last equation above: The f function is XORed twice to L15 with identical inputs (namely R15 and k16). These two outputs of the
92 3 The Data Encryption Standard (DES) and Alternatives f function cancel each other out, so that Rd 1 = L15. Hence, after the first decryption round, we in fact have computed the same values we had before the last encryption round. Thus, the first decryption round reverses the last encryption round. This is an iterative process, which continues in the next 15 decryption rounds and can be expressed as: Ld i = R16−i, Rd i = L16−i, where i = 0, 1, . . . , 16. In particular, after the last decryption round: Ld 16 = R16−16 = R0 Rd 16 = L0 Finally, at the end of the decryption process, we have to reverse the initial per- mutation: IP−1(Rd 16, Ld 16) = IP−1(L0, R0) = IP−1(IP(x)) = x where x is the plaintext that was the input to the DES encryption. 3.5 Security of DES As we discussed in Section 1.2.2, ciphers can be attacked in several ways. With respect to cryptographic attacks, we distinguish between exhaustive key search (or brute-force) attacks and analytical attacks. The latter was demonstrated with the LFSR attack in Section 2.3.2, where we could easily break a stream cipher by solv- ing a system of linear equations. Shortly after DES was proposed, two major criti- cisms against the cryptographic strength of DES centered around two arguments: 1. The key space is too small, i.e., the algorithm is vulnerable against brute-force attacks. 2. The design criteria of the S-boxes were kept secret and there might be an analyti- cal attack that exploits mathematical properties of the S-boxes and which is only known to the DES designers. We discuss both types of attacks below and state the main conclusion about DES security already here: Despite very intensive cryptanalysis since the mid-1970s, cur- rent analytical attacks are not very efficient. However, DES can relatively easily be broken with an exhaustive key-search attack and, thus, plain DES is not suited for most applications anymore.
3.5 Security of DES 93 3.5.1 Exhaustive Key Search The first criticism is nowadays certainly justified. The original cipher proposed by IBM had a key length of 128 bits and it is suspicious that it was reduced to 56 bits. The official statement that a cipher with a shorter key length made it easier to im- plement the DES algorithm on a single chip in 1974 does not sound too convincing. For clarification, let’s recall the principle of an exhaustive key-search (or brute-force attack). Definition 3.5.1 DES exhaustive key search Input: at least one pair of plaintext–ciphertext (x, y) Output: k, such that y = DESk(x) Attack: Test all 256 possible keys until the following condition is fulfilled: DES−1 ki (y) ? = x , i = 0, 1, . . . , 256 − 1 Note that there is a small chance of 1/28 that an incorrect key is found, i.e., a key k which decrypts only the one ciphertext y correctly but not subsequent ciphertexts. If one wants to rule out this possibility, an attacker must check such a key candidate with a second plaintext–ciphertext pair. More about this is found in Section 5.2. Regular computers are not particularly well suited to perform the 256 key tests necessary, but special-purpose key-search machines are an option. It seems highly likely that large (government) institutions have long been able to build such brute- force crackers, which can break DES in a matter of days. In 1977, Whitfield Diffie and Martin Hellman [94] estimated that it was possible to build an exhaustive key- search machine for approximately $20,000,000. Even though they later stated that their cost estimate had been too optimistic, it was clear from the beginning that a cracker could be built with sufficient funding. At the rump session of the CRYPTO 1993 conference, Michael Wiener proposed the design of a very efficient key-search machine which used pipelining techniques. He estimated the cost of his design at approximately $1,000,000, and the time re- quired to find the key at 1.5 days. This was a proposal only, and the machine was not built. In 1998, however, the EFF (Electronic Frontier Foundation) built the hard- ware machine Deep Crack, which performed a brute-force attack against DES in 56 hours. Figure 3.15 shows a photo of Deep Crack. The machine consisted of 1800 in- tegrated circuits, where each had 24 key-test units. The average search time of Deep Crack was 15 days, and the machine was built for less than $250,000. The success- ful break with Deep Crack was considered the official demonstration that DES is no longer secure against determined attacks by many people. Please note that this break does not imply that a weak algorithm had been in use for more than 20 years. It was only possible to build Deep Crack at such a relatively low price because digi- tal hardware had become cheap. In the 1980s it would have been impossible to build a DES cracker without spending many millions of dollars. It can be speculated that
94 3 The Data Encryption Standard (DES) and Alternatives only government agencies were willing to invest such an amount of money for code breaking. Fig. 3.15 Deep Crack — the hardware exhaustive key-search machine that broke DES in 1998 (reproduced with permission from Paul Kocher) DES brute-force attacks also provide an excellent case study for the continuing decrease in hardware costs. In 2006, the COPACOBANA (Cost-Optimized Parallel Code-Breaker) machine was built based on commercial integrated circuits by a team of researchers from the Universities of Bochum and Kiel in Germany (all three au- thors of this book were heavily involved in this effort). COPACOBANA allows one to break DES with an average search time of less than seven days. The interesting part of this undertaking is that the machine could be built with hardware costs in the $10,000 range. Figure 3.16 shows a picture of COPACOBANA. Fig. 3.16 COPACOBANA — A cost-optimized parallel code breaker In summary, a key size of 56 bits is too short to encrypt confidential data nowa- days. Hence, single DES should not be used anymore. However, 3DES, i.e., apply- ing DES three times in a row, yields a much more secure cipher, cf. Section 3.7.2.
3.5 Security of DES 95 Please note that 3DES is currently being phased out by NIST and will officially be discontinued as a U.S. standard after 2023. 3.5.2 Analytical Attacks As was shown in the first chapter, analytical attacks can be very powerful. Since the introduction of DES in the mid-1970s, many excellent researchers in academia (and without doubt many excellent researchers in intelligence agencies) tried to find weaknesses in the structure of DES, which would allow breaking of the cipher. It is a major triumph for the designers of DES that no weakness was found until 1990. In this year, Eli Biham and Adi Shamir discovered what is called differential cryptanalysis (DC). This is a powerful attack, which is in principle applicable to any block cipher. However, it turned out that the DES S-boxes are particularly resistant against this attack. In fact, one member of the original IBM design team declared after the discovery of DC that they had been aware of the attack at the time of design. Allegedly, the reason why the S-box design criteria were not made public was that the design team did not want to make such a powerful attack public. If this claim is true — and all circumstances support it — it means that the IBM and NSA team was 15 years ahead of the research community. It should be noted, however, that in the 1970s and 1980s relatively few people did active research in cryptography. In 1993, a related but distinct analytical attack was published by Mitsuru Matsui, which was named linear cryptanalysis (LC). Similarly to differential cryptanalysis, the effectiveness of this attack heavily depends on the structure of the S-boxes. What is the practical relevance of these two analytical attacks against DES? It turns out that an attacker needs 247 plaintext–ciphertext pairs for a successful DC attack. This assumes particularly chosen plaintext blocks; for random plaintext 255 pairs are needed! In the case of LC, an attacker needs 243 plaintext–ciphertext pairs. All these numbers seem highly impractical for several reasons. First, an attacker needs to know an extremely large number of plaintexts, i.e., pieces of data which are supposedly encrypted and thus hidden from the attacker. Second, collecting and storing such an amount of data takes a long time and requires considerable memory resources. Third, the attack only recovers one key. (This is actually one of many arguments for introducing key freshness in cryptographic applications.) As a result of all these arguments, it does not seem likely that DES can be broken with either DC or LC in real-world systems. However, both DC and LC are very powerful attacks which are applicable to many other block ciphers. Table 3.15 provides an overview of proposed and realized attacks against DES since its standardization. Some entries refer to what is known as the DES Challenges. Starting in 1997, several DES-breaking challenges were organized by the company RSA Security.
96 3 The Data Encryption Standard (DES) and Alternatives Table 3.15 History of full-round DES attacks Date Proposed or implemented attacks 1977 W. Diffie and M. Hellman propose cost estimate for key-search machine 1990 E. Biham and A. Shamir propose differential cryptanalysis, which requires 247 chosen plaintexts 1993 M. Wiener proposes detailed hardware design for key-search machine with an average search time of 36 h and estimated cost of $1,000,000 1993 M. Matsui proposes linear cryptanalysis, which requires 243 chosen ciphertexts Jun. 1997 DES Challenge I broken through brute-force; distributed effort on the internet took 4.5 months Feb. 1998 DES Challenge II–1 broken through brute-force; distributed effort on the inter- net took 39 days Jul. 1998 DES Challenge II–2 broken through brute-force; Electronic Frontier Founda- tion built the Deep Crack key-search machine for about $250,000. The attack took 56 h (15 days average) Jan. 1999 DES Challenge III broken through brute-force by distributed internet effort combined with Deep Crack and a total search time of 22 hours Apr. 2006 Universities of Bochum and Kiel (both in Germany) built the COPACOBANA key-search machine based on low-cost FPGAs for approximately $10,000. Av- erage search time is 7 days. 3.6 Implementation in Software and Hardware In the following, we provide a brief assessment of DES implementation properties in software and hardware. When we talk about software, we refer to DES implemen- tations running on desktop CPUs or embedded microprocessors like smart cards or IoT devices. Hardware refers to DES implementations running on ICs such as application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). Software A straightforward software implementation that follows the data flow of most DES descriptions, such as presented in this chapter, results in a very poor performance. This is due to the fact that many of the atomic DES operations involve bit permuta- tions, in particular the E and P permutations, which are slow in software. Similarly, small S-boxes such as used in DES are efficient in hardware but only moderately efficient on modern CPUs. There have been numerous methods proposed for ac- celerating DES software implementations. The general idea is to use tables with precomputed values of several DES operations, e.g., of several S-boxes and the per- mutation. Optimized implementations require about 240 cycles for encrypting one block on a 32-bit CPU. On a 2-GHz CPU this translates into a theoretical throughput of about 533 Mbits/s. 3DES, which is considerably more secure than single DES,
3.7 DES Alternatives 97 runs at one third of the DES speed. Note that non-optimized implementations are considerably slower, often below 100 Mbit/s. A notable method for accelerating software implementations of DES is bit- slicing, developed by Eli Biham in 1997. The limitation of bit-slicing, however, is that several blocks are encrypted in parallel, which can be a drawback for certain modes of operation such as Cipher Block Chaining (CBC) and Output Feedback (OFB) mode (cf. Chapter 5). Hardware One design criterion for DES was its efficiency in hardware. Permutations such as E, P, IP and IP−1 are very easy to implement in hardware, as they only require wiring but no logic. The small 6-by-4 S-boxes are also relatively easily realizable in hardware. Typically, they are implemented with Boolean logic, i.e., logic gates. On average, one S-box requires about 100 gates. An area-efficient implementation of a single DES round can be done with fewer than 3000 gates. If a high throughput is desired, DES can be made to execute ex- tremely quickly by fitting multiple rounds in one circuit, e.g., by using pipelining. On modern ASICs and FPGAs throughput rates of several 100 Gbit/sec are possi- ble. At the other end of the performance spectrum, very small implementations with fewer than 3000 gates even fit onto low-cost radio frequency identification (RFID) chips. 3.7 DES Alternatives There exist hundreds of other block ciphers. Even though many proposed ciphers have have security weaknesses or have not been well investigated, there are also many block ciphers that are believed to be very secure. In the following a brief list of DES alternatives is discussed. 3.7.1 The Advanced Encryption Standard (AES) and the AES Finalist Ciphers Today, the algorithm of choice for many, many applications has become the Ad- vanced Encryption Standard (AES), which will be introduced in detail in the fol- lowing chapter. With its three key lengths of 128, 192 and 256 bits, AES will be secure against brute-force attacks for several decades, and no analytical attacks with any reasonable chance of success are known. Please note that AES-128 can poten- tially be broken with quantum computers, should they become available in the fu-
98 3 The Data Encryption Standard (DES) and Alternatives ture. However, AES-192 and AES-256 are believed to withstand quantum computer attacks too. Section 12.1 provides more information about this issue. AES was the result of an open competition, and in the last stage of the selection process there were four other finalist algorithms. These are the block ciphers Mars, RC6, Serpent and Twofish. All of them are cryptographically strong and quite fast, especially in software. Based on today’s knowledge, they can all be recommended. The designers of Mars, Serpent and Twofish have always allowed royalty-free use of their ciphers. The patent for RC6 has also expired by now. 3.7.2 Triple DES (3DES) and DESX Triple DES, also denoted 3DES, TDES, Triple DEA or TDEA, has been widely used since the 1990s. 3DES is in particular popular for applications in the payment indus- try. However, 3DES will be phased out for U.S. government applications in 2023. 3DES consists of three subsequent DES operations encrytion-decryption-en- cryption (3DES-EDE): y = DESk3 (DES−1 k2 (DESk1 (x))) as shown in Figure 3.17. The reason for using 3DES in the EDE mode is that it performs a single DES encryption if k3 = k2 = k1, which was desirable for imple- mentations that should also support single DES for legacy reasons. There are two ways to select the three keys. The method preferred today (and more natural) is to choose three independent keys k1, k2 and k3, referred to as 3TDEA. The second method is to choose k2 unique and keys k1 = k3, referred to as 2TDEA. In some older applications it was considered advantageous that only 112 bits of key material was needed for 2TDEA, as opposed to 168 bits for 3TDEA. However, the use of 2TDEA is not encouraged anymore.-1 Fig. 3.17 Triple DES (3DES-EDE)
3.7 DES Alternatives 99 3DES is resistant against brute-force attacks and most analytical attacks. There are subtle security weaknesses when large amounts of data are encrypted under the same key. To avoid such attacks, NIST mandates that not more than 220 blocks of plaintext are encrypted with 3TDEA under the same key. With respect to implemen- tation properties, 3DES is very efficient in hardware but not particularly in software. One might wonder why there is no version of DES that uses double-encryption (“2DES”). This has to do with the meet-in-the-middle attack, which is discussed in Section 5.3.1. In short, 2DES is — surprisingly — not considerably more secure than single DES. DESX is a different approach for strengthening DES by using key whitening. For this, two additional 64-bit keys k1 and k2 are XORed to the plaintext and cipher- text, respectively, prior to and after the DES algorithm. This yields the following encryption scheme: y = DESk,k1,k2 (x) = DESk(x ⊕ k1) ⊕ k2 This surprisingly simple modification makes DES much more resistant against ex- haustive key searches. More about key whitening is said in Section 5.3.3. 3.7.3 Lightweight Cipher PRESENT Since about 2007, several new encryption algorithms that are classified as “light- weight ciphers” have been proposed. Lightweight commonly refers to algorithms with a very low implementation complexity, especially in hardware. Trivium (Sec- tion 2.4.3) is an example of a lightweight stream cipher. A promising block cipher candidate is PRESENT, which was designed specifically for applications such as RFID tags or other internet-of-things (IoT) devices, which are extremely power or cost constrained. (One of the book authors participated in the design of PRESENT.) With respect to other lightweight ciphers and NIST’s standardization efforts in this area, we refer to Section 3.8. Unlike DES, PRESENT is not based on a Feistel network. Instead it is a substitution-permutation network (SP-network) and consists of 31 rounds. We note that AES is also based on an SP-network. The block length is 64 bits, and two key lengths of 80 and 128 bits are supported. A block diagram of the cipher is shown in Figure 3.18. Each of the 31 rounds consists of an XOR operation to introduce a round key Ki, a nonlinear substitution layer (sBoxLayer) and a linear bitwise per- mutation (pLayer). After the last round, the final subkey k32 is applied to the data path. The nonlinear layer uses a 4-bit S-box S, which is applied 16 times in parallel in each round. The key schedule generates 32 round keys from the user-supplied key. Here is the pseudo code:
100 3 The Data Encryption Standard (DES) and Alternativesplaintext K1 K32 sBoxLayer pLayer sBoxLayer pLayer cipher key update update . . . . . . Fig. 3.18 Internal structure of the block cipher PRESENT Pseudo code of the block cipher PRESENT 1 generateRoundKeys() 2 FOR i = 1 TO 31 2.1 addRoundKey(STATE,Ki) 2.2 sBoxLayer(STATE) 2.3 pLayer(STATE) 3 addRoundKey(STATE,K32) We discuss the details of the three steps of PRESENT below. Again, we recom- mend to also look at the diagram in Figure 3.18. addRoundKey At the beginning of each round, the round key Ki is XORed to the current STATE. sBoxLayer PRESENT uses a single 4-bit to 4-bit S-box. This is a direct conse- quence of the pursuit of hardware efficiency, since such an S-box allows a much more compact implementation than, e.g., an 8-bit S-box. The S-box entries in hex- adecimal notation are given in Table 3.16.
3.7 DES Alternatives 101 Table 3.16 The PRESENT S-box in hexadecimal notation x 0 1 2 3 4 5 6 7 8 9 A B C D E F S[x] C 5 6 B 9 0 A D 3 E F 8 4 7 1 2 The 64-bit data path (b63 . . . b0) is referred to as the state. For the sBoxLayer, it is helpful to view the state as consisting of sixteen 4-bit nibbles (w15 . . . w0), where wi = b4·i+3||b4·i+2|| b4·i+1||b4·i for 0 ≤ i ≤ 15, and the output consists of the sixteen nibbles S[wi]. pLayer As in DES, the mixing layer was chosen as a bit permutation, which can be implemented extremely compactly in hardware. The bit permutation used in PRESENT is given by Table 3.17. Bit i of the input state is moved to bit position P(i). Table 3.17 The permutation layer of PRESENT i 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 P(i) 0 16 32 48 1 17 33 49 2 18 34 50 3 19 35 51 i 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 P(i) 4 20 36 52 5 21 37 53 6 22 38 54 7 23 39 55 i 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 P(i) 8 24 40 56 9 25 41 57 10 26 42 58 11 27 43 59 i 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 P(i) 12 28 44 60 13 29 45 61 14 30 46 62 15 31 47 63 The bit permutation is quite regular and can in fact be expressed as: P(i) = { i · 16 mod 63, i = 0, . . . , 62 63, i = 63 Key Schedule of PRESENT-80 First, we describe the key schedule for PRESENT with an 80-bit key. This key length is attractive for applications that have only short- term security requirements, e.g., low-cost IoT devices. For all other applications the 128-bit key is recommended, which is described below. The user-supplied key is stored in a key register K and is represented by the 80 bits k79k78 . . . k0. At round i the 64-bit round key Ki consists of the 64 leftmost bits of the current contents of register K. Thus at round i we have: Ki = k79k78 . . . k16
102 3 The Data Encryption Standard (DES) and Alternatives The first subkey K1 is a direct copy of the 64 leftmost bits of the user-supplied key. For the following subkeys K2, . . . , K32 the key register K = k79k78 . . . k0 is updated as follows: Step 1: [k79k78 . . . k1k0] = [k18k17 . . . k20k19] Step 2: [k79k78k77k76] = S[k79k78k77k76] Step 3: [k19k18k17k16k15] = [k19k18k17k16k15] ⊕ round_counter Here is an explanation of the three steps of the key schedule: Step 1: The key register is rotated by 61 bit positions to the left. Step 2: The leftmost four bits are passed through the PRESENT S-box. Step 3: The round_counter value i is XORed with bits k19k18k17k16k15 of K, where the least significant bit of round_counter is on the right. This counter is a simple integer which takes the values (00001, 00010, . . . , 11111). For example, for the derivation of K2 the counter value 00001 is used; for K3, the counter value 00010; and so on. Key Schedule of PRESENT-128 The key is stored in a register K and represented by the bits k127k126 . . . k0. In round i, the 64-bit round key Ki consists of the 64 leftmost bits of the current contents of register K, i.e., Ki = k127k126 . . . k64 The first subkey K1 is a direct copy of the 64 leftmost bit of the user-supplied key. For the following subkeys K2, . . . , K32, the key register K = k127k126 . . . k0 is updated as follows: Step 1: [k127k126 . . . k1k0] = [k66k65 . . . k68k67] Step 2: [k127k126k125k124] = S[k127k126k125k124] Step 3: [k123k122k121k120] = S[k123k122k121k120] Step 4: [k66k65k64k63k62] = [k66k65k64k63k62] ⊕ round_counter Implementation and Security PRESENT-80 can be implemented in hardware with an area of approximately 1600 gate equivalences, where the encryption of one 64-bit plaintext block requires 32 clock cycles. As an example, even at a relatively slow clock rate of 1 MHz, which is quite typical on low-cost, low-energy devices, a remarkably high throughput of 2 Mbit/s is achieved. It is possible to realize the cipher with as few as approximately 1000 gate equivalences, where the encryption of one 64-bit plaintext requires 547 clock cycles. A fully pipelined implementation of PRESENT with 31 encryption stages achieves a throughput of 64 bits per clock cycle, which can result in encryption throughputs of more than 50 Gbit/s. As a result of the aggressively hardware-optimized design of PRESENT, its soft- ware performance is slower compared to many other modern ciphers like AES. An optimized software implementation on a Pentium III CPU in C achieves a through- put of about 60 Mbit/s at a frequency of 1 GHz. However, it performs quite well on small microprocessors, which are common in inexpensive consumer products.
3.8 Discussion and Further Reading 103 At the time of writing, no attacks are known against the full-round version of PRESENT that are better than a brute-force attacks. 3.8 Discussion and Further Reading DES History and Attacks Even though single DES (i.e., non-3DES) is hardly used anymore, its history helps us understand the evolution of cryptography since the mid-1970s from an obscure discipline almost solely studied in government or- ganizations towards an open field with many players in industry and academia. A summary of the history of DES can be found in [245]. Today, the two major analyt- ical attacks developed against DES, differential and linear cryptanalysis, are among the most powerful general methods for breaking block ciphers. For readers inter- ested in the theory of block ciphers, including differential and linear cryptanalysis, the very accessible book [161] is recommended. The original references for differ- ential and linear cryptanalysis are [48, 181]. As we have seen in this chapter, DES should no longer be used since a brute-force attack can be accomplished at low cost in little time with cryptanalytical hardware. The two machines built outside governments, Deep Crack and COPACOBANA, are instructive examples of how to build low-cost “supercomputers” for very narrowly defined computational tasks. More information about Deep Crack can be found in [118] and about COPACOBANA in the articles [166, 137] and online at [76]. Earlier, Michael Wiener proposed (but did not built) a very efficient key-search ma- chine which used pipelining techniques. An update of his proposal can be found in [252]. Readers interested in the fascinating area of cryptanalytical computers in general should take a look at the SHARCS (Special-purpose Hardware for Attacking Cryptographic Systems) workshop series, which took place irregularly from 2005 until 2012 [249]. In July 2017, NIST first mentioned the retiring of 3DES [83], following a new security analysis that is based on so-called collision attacks. It is described in Ref- erence [45]. The attack can be mitigated by limiting the number of plaintext blocks which are encrypted under one key. In November 2017, NIST restricted the usage of 3DES to 220 64-bit blocks of plaintext (8 MB of data) using a single set of keys. As a consequence, it should no longer be used for TLS, IPsec or large file encryption applications [26]. At the time of writing, 3DES (or TDEA) is being phased out by NIST and will officially be discontinued as a U.S. standard after 2023. A guideline for the transitioning away from 3DES is provided in [27]. DES Implementation With respect to software implementation of DES, an early reference is [47]. More advanced techniques are described in [167]. The powerful method of bit-slicing, which we described in Section 3.6, is applicable not only to DES but to most other ciphers. Regarding DES hardware implementation, an early but still very interesting ref- erence is [247]. There are many descriptions of high-performance implementations
104 3 The Data Encryption Standard (DES) and Alternatives of DES on a variety of hardware platforms, including FPGAs [244], standard ASICs as well as more exotic semiconductor technology [107]. DES Alternatives and Lightweight Ciphers It should be noted that hundreds of block ciphers have been proposed since DES came into existence in the mid-1970s. DES has influenced the design of many other encryption algorithms, especially those proposed in the 1980s and 1990s. Some of the most popular block ciphers are also based on Feistel networks. Examples of other Feistel ciphers from this era include Blowfish, CAST, KASUMI, Mars, MISTY1, Twofish and RC6. One cipher from the pre-2000 area that is well known and markedly different from DES is IDEA; it uses arithmetic in three different algebraic structures as atomic operations. All this said, the block cipher of choice for many of today’s applications is AES, which is introduced in the following chapter. DES is a good example of a block cipher which is very efficient in hardware. There are applications that are extremely cost sensitive and power constrained, e.g., RFID tags or other low-cost IoT devices, for which such lightweight ciphers are very attractive. Good references for PRESENT are [59, 221]. In addition to PRESENT, other proposed very small block ciphers include Clefia [77], HIGHT [144] and mCrypton [175]. PRESENT and Clefia have been standardized in ISO/IEC 29192- 2 [150]. In 2015, NIST started the process of standardizing lightweight ciphers. In 2019, 57 candidate algorithms were submitted. The subsequent selection process stretched over three rounds. After Round 2, ten finalist algorithms were announced in early 2021. In February 2023, the Ascon family of lightweight ciphers was selected as a future NIST standard. The algorithm was designed by a team of European cryptog- raphers and a description of the cipher is given in [98]. Interestingly, Ascon uses neither a Feistel construction (like DES) nor a substitution-permutation network (like AES and PRESENT) but rather a sponge construction. More about sponge constructions will be said in Section 11.5.1 in the context of the SHA-3 hash func- tion. Among the many proposals for lightweight ciphers, the algorithms Simon and Speck [31] play a particular role. They are efficient when implemented in software and hardware but they are probably best known for the fact that they are designed by the NSA. There have been controversies around efforts to include Simon and Speck in industrial standards by ISO [19], as some countries were worried about possible weaknesses in the ciphers.
3.9 Lessons Learned 105 3.9 Lessons Learned DES was the dominant symmetric encryption algorithm from the mid-1970s to the mid-1990s. Since ciphers with 56-bit keys were no longer secure, the Ad- vanced Encryption Standard (AES) was created as DES’s replacement. Standard DES can be broken relatively easily nowadays through an exhaustive key search due its short key length of 56 bits. DES is quite robust against known analytical attacks: In practice it is very diffi- cult to break the cipher with differential or linear cryptanalysis. DES is reasonably efficient in software but fast and small in hardware. By encrypting with DES three times in a row, triple DES (3DES) is created, which is still secure if the amount of data encrypted under one set of keys is limited. The “default” symmetric cipher nowadays is often AES. In addition, the other four AES finalist ciphers all seem very secure and efficient. Since about 2005 several proposals for lightweight ciphers have been made. They are suited for resource-constrained applications.
106 3 The Data Encryption Standard (DES) and Alternatives Problems 3.1. As stated in Section 3.5.2, one important property that ensures that DES is secure is that the S-boxes are nonlinear. In this problem we verify this property by computing the output of S1 for several pairs of inputs. Show that S1(x1) ⊕ S1(x2) 6 = S1(x1 ⊕ x2), where “⊕” denotes bitwise XOR, for: 1. x1 = 000000, x2 = 000001 2. x1 = 111111, x2 = 100000 3. x1 = 101010, x2 = 010101 3.2. The S-box S4 has special properties: 1. Show that the 1st row can be computed from the 0th row with the help of the following mapping: (y1, y2, y3, y4) → (y2, y1, y4, y3) ⊕ (0, 1, 1, 0) where (y1, y2, y3, y4) denotes the binary output of the S-box. It is sufficient to show the mapping for the first five entries. Note that “row” refers to the standard representation of S-boxes, which we also use in this book. 2. Show that the same holds for rows 2 and 3. 3.3. We want to verify that IP(·) and IP−1(·) are truly inverse operations. We con- sider a vector x = (x1, x2, . . . , x64) of 64 bits. Show that IP−1(IP(x)) = x holds for the first five bits of x, i.e., for xi, i = 1, 2, 3, 4, 5. 3.4. What is the output of the first round of the DES algorithm when the plaintext and the key are both all zeros? 3.5. What is the output of the first round of the DES algorithm when the plaintext and the key are both all ones? 3.6. Remember that it is desirable for good block ciphers that a change in one input bit affects many output bits, a property that is called diffusion or the avalanche effect. We try now to get a feeling for the avalanche property of DES. We apply an input word that has a “1” at bit position 57 and all other bits as well as the key are zero. (Note that the input word has to run through the initial permutation.) 1. How many S-boxes get a different input compared to the case when an all-zero plaintext is provided? 2. What is the minimum number of output bits of the S-boxes that will change according to the S-box design criteria? 3. What is the output after the first round?
3.9 Problems 107 4. How many output bits after the first round have actually changed compared to the case when the plaintext is all zero? (Observe that we only consider a single round here. There will be more and more output differences after every new round. Hence the term avalanche effect.) 3.7. An avalanche effect is also desirable for the key: A one-bit change in a key should result in a dramatically different ciphertext if the plaintext is unchanged. 1. Assume an encryption with a given key. Now assume the key bit at position 1 (prior to PC–1) is flipped. Which S-boxes in which rounds are affected by the bit flip during DES encryption? 2. Which S-boxes in which DES rounds are affected by this bit flip during DES decryption? 3.8. In this problem we look at the relationship between the DES round keys and the original key. It turns out that each of the 48 bits of every round key k1, . . . , k16 is a direct map of one bit of the original 64-bit input key k. 1. Determine which of the bits of k form the first two bits of the round key k1. 2. Determine which of the bits of k form the first two bits of the round key k2. 3.9. A DES key Kw is called a weak key if encryption and decryption are identical operations: DESKw (x) = DES−1 Kw (x), for all x (3.1) 1. Describe the relationship of the subkeys in the encryption and decryption algo- rithm that is required so that Equation (3.1) is fulfilled. 2. There are four weak DES keys. What are they? 3. What is the likelihood that a randomly selected key is weak? 3.10. DES has a somewhat surprising property related to bitwise complements of its inputs and outputs. We investigate the property in this problem. We denote the bitwise complement (that is, all bits are flipped) of a number A by A′. Let ⊕ denote bitwise XOR. We want to show that if y = DESk(x) then y′ = DESk′ (x′) (3.2) This states that if we complement the plaintext and the key, then the ciphertext output will also be the complement of the original ciphertext.
108 3 The Data Encryption Standard (DES) and Alternatives Your task is to prove this property. The proof can be done along the following steps: 1. Show that for any bit strings A, B of equal length, A′ ⊕ B′ = A ⊕ B and A′ ⊕ B = (A ⊕ B)′ (These two operations are needed for some of the following steps.) 2. Show that PC–1(k′) = (PC–1(k))′. 3. Show that LSi(C′ i−1) = (LSi(Ci−1))′. 4. Using the two results from above, show that if ki are the subkeys generated from k, then k′ i are the subkeys of k′, where i = 1, 2, . . . , 16. 5. Show that IP(x′) = (IP(x))′. 6. Show that E(R′ i) = (E(Ri))′. 7. Using all previous results, show that if Ri−1, Li−1 and ki generates Ri, then R′ i−1, L′ i−1, and k′ i generates R′ i. 8. Show that Equation (3.2) is true. 3.11. Assume we perform a known-plaintext attack against DES with one pair of plaintext and ciphertext. How many keys do we have to test in a worst-case sce- nario if we apply an exhaustive key search in a straightforward way? How many on average? 3.12. In this problem we want to study the clock frequency requirements for a hard- ware implementation of DES in real-world applications. The speed of a DES im- plementation is mainly determined by the time required to compute one round. The hardware block for one round is used 16 consecutive times in order to generate the encrypted output. (An alternative approach would be to build a hardware pipeline with 16 stages, resulting in 16-fold increased hardware costs but we are not looking at such a pipelined implementation in this problem.) 1. Let’s assume that computing the round function can be performed in one clock cycle. Develop an expression for the required clock frequency for encrypting a stream of data with a data rate r [bits/sec]. Ignore the time needed for the initial and final permutation. 2. What clock frequency is required for encrypting a network link running at a speed of 1 Gb/sec? What is the clock frequency if we want to support a speed of 8 Gb/sec? 3.13. In this example we want to get a feeling for performing a brute-force attack on a 56-bit key. For this purpose, we study the COPACOBANA key-search machine (cf. Section 3.5.1). 1. Compute the run time of an average exhaustive key search on DES assuming the following implementational details:
3.9 Problems 109 We use the COPACOBANA machine with 20 FPGA modules. 6 FPGAs per FPGA module. 4 DES engines per FPGA. Each DES engine is fully pipelined and is capable of performing one encryp- tion per clock cycle. 100 MHz clock frequency. 2. How many COPACOBANA machines do we need if we want to have an average search time of one hour? (We note that COPACOBANA was designed in 2006 and current hardware will be even more powerful.) 3. Why does any design of a key-search machine constitute only an upper secu- rity threshold? By upper security threshold we mean a (complexity) measure that describes the maximum security that is provided by a given cryptographic algorithm. 3.14. In this problem, we study a real-world case of a weak password-based key derivation. A commercial file encryption program from the early 1990s used stan- dard DES with 56 key bits. In those days, performing an exhaustive key search was considerably harder than today, and thus the key length was sufficient for some applications. Unfortunately, the implementation of the key generation was flawed, which we are going to analyze. Assume that we can test 106 keys per second on a conventional PC. The key is generated from a password consisting of 8 characters? The key is a simple concatenation of the 8 ASCII characters, yielding 64 = 8 · 8 key bits. With the permutation PC–1 in the key schedule, the least significant bit (LSB) of each 8-bit character is ignored, yielding 56 key bits. 1. What is the size of the key space if all 8 characters are randomly chosen 8-bit ASCII characters? How long does an average key search take with a single PC? 2. How many key bits are used if the 8 characters are randomly chosen 7-bit ASCII characters (i.e., the most significant bit is always zero)? How long does an aver- age key search take with a single PC? 3. How large is the key space if, in addition to the restriction in Part 2, only let- ters are used as characters. Furthermore, unfortunately, all letters are converted to capital letters before generating the key in the software. How long does an average key search take with a single PC?
110 3 The Data Encryption Standard (DES) and Alternatives 3.15. This problem deals with the lightweight cipher PRESENT. 1. Calculate the state of PRESENT-80 after the execution of one round. We rec- ommend using the table shown below and to solve the problem with paper and pencil. The following values are given (in hexadecimal notation): plaintext = 0000 0000 0000 0000 key = BBBB 5555 5555 EEEE FFFF. Plaintext 0000 0000 0000 0000 Round key State after KeyAdd State after sBoxLayer State after pLayer 2. Now calculate the round key for the second round using the following table. Key BBBB 5555 5555 EEEE FFFF Key state after rotation Key state after sBoxLayer Key state after CounterAdd Round key for Round 2
Chapter 4 The Advanced Encryption Standard (AES) The Advanced Encryption Standard (AES) is the most widely used symmetric ci- pher today. Even though the term “Standard” in its name originally only referred to U.S. government applications, the AES block cipher has been adopted by many industry standards and is used in numerous commercial systems. Examples of stan- dards that incorporate AES are the web security protocol TLS, the internet security standard IPsec and the Wi-Fi encryption standard IEEE 802.11i. Countless other applications, such as the instant messenger WhatsApp, password managers and file encryption software make use of the block cipher too. To date, no attacks against AES significantly better than brute-force are known. In this chapter, you will learn: The design process of the U.S. symmetric encryption standard, AES The encryption and decryption function of AES The internal structure of AES, namely: byte substitution layer diffusion layer key addition layer key schedule Basic facts about Galois fields Efficiency of AES implementations 111 C. Paar et al., Understanding Cryptography, https://doi.org/10.1007/978-3-662-69007-9_4 © The Editor(s) (if applicable) and The Author(s), under exclusive license to Springer-Verlag GmbH, DE, part of Springer Nature 2024
112 4 The Advanced Encryption Standard (AES) 4.1 Introduction In 1999, the U.S. National Institute of Standards and Technology (NIST) indicated that DES should only be used for legacy systems, and instead triple DES (3DES) should be used. Even though 3DES resists brute-force attacks with today’s tech- nology, there are several problems with it. First, it is not very efficient with regard to software implementations. A more serious drawback is its block size of 64 bits, which gives rise to certain attacks (cf. Section 3.8) if large blocks of data are en- crypted. The short block size also makes it more difficult to build a hash function from 3DES (cf. Section 11.3.1). Finally, if one is worried about attacks with quan- tum computers in the future, key lengths on the order of 256 bits are desirable. All these considerations led NIST to the conclusion that an entirely new block cipher was needed as a replacement for DES. In 1997, NIST called for proposals for a new Advanced Encryption Standard. Unlike the development of DES, the selection of the AES algorithm was an open process administered by NIST. In three subsequent evaluation rounds, NIST and the international scientific community discussed the advantages and disadvantages of the submitted ciphers and narrowed down the number of potential candidates. In 2001, NIST declared the cipher Rijndael as the new AES and published it as a U.S. standard (FIPS PUB 197). Rijndael was designed by two young Belgian cryptographers. Within the call for proposals, the following requirements for all AES candidate submissions were mandatory: block cipher with 128-bit block size, three key lengths must be supported: 128, 192 and 256 bits, security relative to other submitted algorithms, efficiency in software and hardware. The invitation for submitting suitable algorithms and the subsequent evaluation of the successor of DES was a public process. A compact chronology of the AES selection process is given here: The need for a new block cipher was announced in January 1997 by NIST. A formal call for AES was announced in September 1997. Fifteen candidate algorithms were submitted by researchers from several coun- tries by August 1998. In August 1999, five finalist algorithms were announced: Mars by IBM Corporation, RC6 by RSA Laboratories, Rijndael, by Joan Daemen and Vincent Rijmen, Serpent, by Ross Anderson, Eli Biham and Lars Knudsen, Twofish, by Bruce Schneier, John Kelsey, Doug Whiting, David Wagner, Chris Hall and Niels Ferguson. On October 2, 2000, NIST announced that it had chosen Rijndael as the AES.
4.2 Overview of the AES Algorithm 113 On November 26, 2001, AES was formally approved as a U.S. federal standard. It is expected that AES will stay the dominant symmetric-key algorithm for many commercial applications for the next few decades, especially in the Western world. It is also remarkable that in 2003, the U.S. National Security Agency (NSA) an- nounced that it allows AES to encrypt classified documents up to the level SECRET for all key lengths and up to the TOP SECRET level for key lengths of either 192 or 256 bits. Prior to that date, only non-public algorithms had been used for the encryption of classified documents. 4.2 Overview of the AES Algorithm The AES cipher is almost identical to the block cipher Rijndael. The Rijndael block and key size vary between 128, 192 and 256 bits. However, the AES standard only calls for a block size of 128 bits. Hence, only Rijndael with a block length of 128 bits is known as the AES algorithm. In the remainder of this chapter, we only discuss the standardized version of Rijndael with a block size of 128 bits, cf. Figure 4.1.AES k 128/192/256 128 128 x y Fig. 4.1 AES input and output parameters The three key lengths supported by AES were a NIST design requirement. The number of internal rounds of the cipher is a function of the key length, according to Table 4.1. Table 4.1 Key lengths and number of rounds for AES key length # rounds = nr 128 bits 10 192 bits 12 256 bits 14
114 4 The Advanced Encryption Standard (AES) In contrast to DES, AES does not have a Feistel structure. Feistel networks do not encrypt an entire block per iteration, e.g., in DES, 64/2 = 32 bits are encrypted in one round. AES, on the other hand, encrypts all 128 bits in one iteration. This is one reason why it has a comparably small number of rounds. Each AES round consists of so-called layers. There are only three different types of layers. Each layer manipulates all 128 bits of the data path. The data path is also referred to as the state of the algorithm. Each round, with the exception of the last, consists of all layers as shown in Figure 4.2: the plaintext is denoted by x, the ciphertext by y and the number of rounds by nr. Moreover, the last round nr does not make use of the MixColumn transformation, which makes the encryption and decryption scheme symmetric. We continue with a brief description of the layers: Key addition layer A 128-bit round key, or subkey, which has been derived from the main key in the key schedule, is XORed to the state. Byte substitution layer (S-box) Each element of the state is nonlinearly trans- formed using lookup tables with special mathematical properties. This introduces confusion to the data, i.e., it ensures that changes in individual bits lead to nonlinear changes in the state. Diffusion layer This provides diffusion to the state, i.e., it ensures that changes of individual bits propagate quickly across the 128 bits of the data path. It consists of two sublayers, both of which perform linear operations: The ShiftRows sublayer permutes the data on a byte level. The MixColumn sublayer is a matrix operation that combines (or mixes) blocks of four bytes. The key schedule computes the round keys, or subkeys, (k0, k1, . . . , knr ) from the user-provided AES key. We note that there are nr + 1 subkeys, e.g., for the (widely used) 10-round version of AES, there are 11 subkeys. Before we describe the internal functions of the layers in Section 4.4, we have to introduce a new mathematical concept, namely Galois fields. 4.3 Some Mathematics: A Brief Introduction to Galois Fields In AES, Galois field arithmetic is used in most layers, especially in the S-box and the MixColumn layer. Hence, for a deeper understanding of the internals of AES, we provide an introduction to Galois fields as needed for this purpose before we continue with the actual cipher description in Section 4.4. A background in Galois fields is not required for a basic understanding of AES, and the reader interested in this can skip this section.
4.3 Some Mathematics: A Brief Introduction to Galois Fields 115Transform Byte Substitution Layer Key Addition Layer MixColumn Layer ShiftRows Layer Diffusion Layer Key Addition Layer x Plaintext Byte Substitution Layer Key Addition Layer MixColumn Layer ShiftRows Layer Key Addition Layer ShiftRows Layer Byte Substitution Layer xy=AES( ) rkn k rn −1 last round round −1 round 1 nr rn k r r Key 0k 1k Ciphertext 1Transform 0Transform n −1Transform n Fig. 4.2 AES encryption block diagram
116 4 The Advanced Encryption Standard (AES) 4.3.1 Existence of Finite Fields A finite field, sometimes also called a Galois field, is a set with a finite number of elements. Roughly speaking, a Galois field is a finite set of elements in which we can add, subtract, multiply and invert. Before we introduce the definition of a field, we first need the concept of a simpler algebraic structure, a group. Definition 4.3.1 Group A group is a set of elements G together with an operation ◦ that combines two elements of G. A group has the following properties: 1. The group operation ◦ is closed. That is, for all a, b ∈ G, it holds that a ◦ b = c ∈ G. 2. The group operation is associative. That is, a◦(b◦c) = (a◦b)◦c for all a, b, c ∈ G. 3. There is an element 1 ∈ G, called the neutral element (or identity element), such that a ◦ 1 = 1 ◦ a = a for all a ∈ G. 4. For each a ∈ G there exists an element a−1 ∈ G, called the in- verse of a, such that a ◦ a−1 = a−1 ◦ a = 1. 5. A group G is abelian (or commutative) if, furthermore, a ◦ b = b ◦ a for all a, b ∈ G. A group is a set with one operation and the corresponding inverse operation. If the operation is called addition, the inverse operation is subtraction; if the operation is multiplication, the inverse operation is division (or multiplication with the inverse element). Example 4.1. The set of integers Zm = {0, 1, . . . , m − 1} and the operation addition modulo m form a group with the neutral element 0. Every element a has an inverse −a such that a + (−a) ≡ 0 mod m. Note that this set does not form a group with the operation multiplication because not all elements a have an inverse such that a a−1 ≡ 1 mod m. In order to have all four basic arithmetic operations (i.e., addition, subtraction, multiplication, division), we need a set that contains an additive and a multiplicative abelian group. This is what we call a field.
4.3 Some Mathematics: A Brief Introduction to Galois Fields 117 Definition 4.3.2 Field A field F is a set of elements with the following properties: All elements of F form an additive abelian group with the group operation “+” and the neutral element 0. All elements of F except 0 form a multiplicative abelian group with the group operation “×” and the neutral element 1. When the two group operations are mixed, the distributivity law holds, i.e., for all a, b, c ∈ F: a(b + c) = (ab) + (ac). Example 4.2. The set R of real numbers is a field with the neutral element 0 for the additive group and the neutral element 1 for the multiplicative group. Every real number a has an additive inverse, namely −a, and every nonzero element a has a multiplicative inverse a−1 = 1/a. In cryptography, we are almost always interested in fields with a finite number of elements, which we call finite fields or Galois fields. The number of elements in the field is called the order or cardinality of the field. Of fundamental importance is the following theorem. Theorem 4.3.1 A field with order q only exists if q is a prime power, i.e., q = pm, for some positive integer m and prime integer p. p is called the characteristic of the finite field. This theorem implies that there are, for instance, finite fields with 11 elements, 81 elements (since 81 = 34) or 256 elements (since 256 = 28, and 2 is a prime). In contrast, there is no finite field with 12 elements since 12 = 22 · 3, and 12 is thus not a prime power. In the literature, the notations F and GF are both used for Galois fields. In this chapter, we will use the latter one, with GF(pm) denoting a field with pm elements. In the remainder of this section, we look at how finite fields can be built, and more importantly for our purpose, how we can do arithmetic in them. 4.3.2 Prime Fields The most intuitive examples of finite fields are those with a prime order, i.e., fields with m = 1. Elements of the field GF(p) can simply be represented by integers 0, 1, . . . , p − 1. The two operations of the field are modular integer addition and integer multiplication modulo p.
118 4 The Advanced Encryption Standard (AES) Theorem 4.3.2 Let p be a prime. The integer ring Zp is denoted by GF(p) and is referred to as a prime field, or as a Galois field, with a prime number of elements. All nonzero elements of GF(p) have an inverse. Arithmetic in GF(p) is done modulo p. This means that if we consider the integer ring Zm — which was introduced in Section 1.4.2 — and m happens to be a prime, Zm is not only a ring but also a finite field. In order to do arithmetic in a prime field, we have to follow the rules for integer rings: Addition and multiplication are done modulo p, the additive inverse −a of any element a is defined by a + (−a) ≡ 0 mod p, and the multiplicative inverse a−1 of any nonzero element a is defined as a · a−1 ≡ 1 mod p. Let’s have a look at an example of a prime field. Example 4.3. We consider the elements of the finite field GF(5), which are in the set {0, 1, 2, 3, 4}. The tables below describe how to add and multiply any two elements, as well as the additive and multiplicative inverse of the field elements. Using these tables, we can perform all calculations in this field without using modular reduction explicitly. addition + 0 1 2 3 4 0 0 1 2 3 4 1 1 2 3 4 0 2 2 3 4 0 1 3 3 4 0 1 2 4 4 0 1 2 3 additive inverse −0 = 0 −1 = 4 −2 = 3 −3 = 2 −4 = 1 multiplication × 0 1 2 3 4 0 0 0 0 0 0 1 0 1 2 3 4 2 0 2 4 1 3 3 0 3 1 4 2 4 0 4 3 2 1 multiplicative inverse 0−1 does not exist 1−1 = 1 2−1 = 3 3−1 = 2 4−1 = 4 A very important prime field is GF(2), which is the smallest finite field that exists. Let’s have a look at the multiplication and addition tables for the field. Example 4.4. We consider the finite field GF(2) and its elements in the set {0, 1}. Arithmetic is simply done modulo 2, yielding the following arithmetic tables:
4.3 Some Mathematics: A Brief Introduction to Galois Fields 119 addition + 0 1 0 0 1 1 1 0 multiplication × 0 1 0 0 0 1 0 1 We saw already in Chapter 2 that modulo 2 addition, i.e., GF(2) addition, is equivalent to the XOR operation. From the example above we learn that GF(2) multiplication is equivalent to the logical AND operation. 4.3.3 Extension Fields GF(2m) AES makes heavily use of the finite field with 256 elements, which is denoted by GF(28). This field was chosen because each of the field elements can be represented by exactly one byte. For the S-box and MixColumn layers, AES treats every byte of the internal data path as an element of the field GF(28) and manipulates the data by performing arithmetic in this finite field. If the order of a finite field is not prime, and 28 is clearly not a prime, the addition and multiplication operation cannot be realized as integer addition and multiplica- tion modulo 28. Such fields with m > 1 are called extension fields. In order to deal with extension fields, we need (1) a different representation for field elements and (2) different rules for performing arithmetic with the elements. We will see in the following that elements of extension fields can be represented as polynomials and that computation in the extension field is achieved by performing a certain type of polynomial arithmetic. In extension fields GF(2m), elements are not represented as integers but as poly- nomials with coefficients in GF(2). The polynomials have a maximum degree of m − 1, so that there are m coefficients in total for every element. In the field GF(28), which is used in AES, each element A ∈ GF(28) is thus represented as: A(x) = a7x7 + · · · + a1x + a0, ai ∈ GF(2) Note that there are exactly 256 = 28 such polynomials. The set of these 256 polyno- mials encodes the elements of the finite field GF(28). It is also important to observe that every polynomial can simply be stored in digital form as an 8-bit vector A = (a7, a6, a5, a4, a3, a2, a1, a0) In particular, we do not have to store the powers x7, x6, etc. It is clear from the bit positions to which power xi each coefficient belongs.
120 4 The Advanced Encryption Standard (AES) 4.3.4 Addition and Subtraction in GF(2m) Let’s now look at addition and subtraction in extension fields. The key addition layer of AES uses addition. It turns out that these operations are straightforward. They are simply achieved by performing standard polynomial addition and subtraction: We merely add or subtract coefficients with equal powers of x. The coefficient additions or subtractions are done in the underlying field GF(2). Definition 4.3.3 Extension field addition and subtraction Let A(x), B(x) ∈ GF(2m). The sum of the two elements is then com- puted according to: C(x) = A(x) + B(x) = m−1 ∑ i=0 cixi, ci ≡ ai + bi mod 2 and the difference is computed according to: C(x) = A(x) − B(x) = m−1 ∑ i=0 cixi, ci ≡ ai − bi ≡ ai + bi mod 2 Note that we perform modulo 2 addition (or subtraction) with the coefficients. Let’s have a look at an example in the field GF(28). Example 4.5. Here is how the sum C(x) = A(x)+B(x) of two elements from GF(28) is computed: A(x) = x7+ x6+ x4+ 1 B(x) = x4+ x2+ 1 C(x) = x7+ x6+ x2 As we saw in Chapter 2, addition and subtraction modulo 2 are the same operation. Moreover, addition modulo 2 is equal to bitwise XOR. Hence, if we computed the difference A(x) − B(x) of the two polynomials from the example above, we would get the same result as for the sum. 4.3.5 Multiplication in GF(2m) Multiplication in GF(28) is the core operation of the MixColumn layer of AES. As a first step, two elements (represented by their polynomials) of a finite field GF(2m) are multiplied using the standard polynomial multiplication rule:
4.3 Some Mathematics: A Brief Introduction to Galois Fields 121 A(x) · B(x) = (am−1xm−1 + · · · + a0) · (bm−1xm−1 + · · · + b0) C′(x) = c′ 2m−2x2m−2 + · · · + c′ 0 where: c′ 0 = a0b0 mod 2 c′ 1 = a0b1 + a1b0 mod 2 ... c′ 2m−2 = am−1bm−1 mod 2 Note that all coefficients ai, bi and ci are elements of GF(2), and that coefficient arithmetic is performed in GF(2). In general, the product polynomial C(x) will have a degree higher than m − 1 and has to be reduced. The basic idea is an approach similar to the case of multiplication in prime fields: In GF(p), we multiply the two integers, divide the result by a prime, and consider only the remainder. Here is what we do in extension fields: The product of the multiplication is divided by a certain polynomial, and we consider only the remainder after the polynomial division. We need irreducible polynomials for the modulo reduction. We recall from Section 2.3.1 that irreducible polynomials are roughly comparable to prime numbers, i.e., their only factors are 1 and the polynomial itself. Definition 4.3.4 Extension field multiplication Let A(x), B(x) ∈ GF(2m) and let P(x) ≡ m ∑ i=0 pixi, pi ∈ GF(2) be an irreducible polynomial of degree m. Multiplication of the two elements A(x), B(x) is performed as C(x) ≡ A(x) · B(x) mod P(x) Thus, every field GF(2m) requires an irreducible polynomial P(x) of degree m with coefficients from GF(2). Note that not all polynomials are irreducible. For example, the polynomial x4 + x3 + x + 1 is reducible since x4 + x3 + x + 1 = (x2 + x + 1)(x2 + 1) and hence cannot be used to construct the extension field GF(24). Since primitive polynomials are a special type of irreducible polynomials, the polynomials in Ta- ble 2.2 can be used for constructing fields GF(2m). For AES, the irreducible poly- nomial P(x) = x8 + x4 + x3 + x + 1
122 4 The Advanced Encryption Standard (AES) is used. It is part of the AES specification. Example 4.6. We want to multiply the two polynomials A(x) = x3 + x2 + 1 and B(x) = x2 + x in the field GF(24). The irreducible polynomial of this Galois field is given as P(x) = x4 + x + 1 The plain polynomial product is computed as: C′(x) = A(x) · B(x) = x5 + x3 + x2 + x We can now reduce C′(x) using the polynomial division method we learned in school. However, sometimes it is easier to reduce each of the leading terms x4 and x5 individually by using the following equivalences: x4 = 1 · P(x) + (x + 1) x4 ≡ x + 1 mod P(x) x5 ≡ x2 + x mod P(x) Now, we only have to insert the reduced expression for x5 into the intermediate result C′(x): C(x) ≡ x5 + x3 + x2 + x mod P(x) C(x) ≡ (x2 + x) + (x3 + x2 + x) = x3 A(x) · B(x) ≡ x3 It is important not to confuse multiplication in GF(2m) with integer multiplica- tion, especially if we are concerned with software implementations of Galois fields. Recall that the polynomials, i.e., the field elements, are normally stored in a com- puter as bit vectors. If we look at the multiplication from the previous example, the following very atypical operation is being performed on the bit level: A · B = C (x3 + x2 + 1) · (x2 + x) = x3 (1 1 0 1) · (0 1 1 0) = (1 0 0 0) This computation is not identical to integer arithmetic. If the polynomials are in- terpreted as integers, i.e., (1101)2 = 1310 and (0110)2 = 610, the result is (1001110)2 = 7810, which is clearly not the same as the Galois field multiplication product. Hence, even though we can represent field elements as integer data types, we cannot make use of the integer arithmetic provided by computers!
4.3 Some Mathematics: A Brief Introduction to Galois Fields 123 4.3.6 Inversion in GF(2m) Inversion in GF(28) is the core operation of the Byte Substitution layer, which contains the AES S-boxes. For a given finite field GF(2m) and the correspond- ing irreducible reduction polynomial P(x), the inverse A−1 of a nonzero element A ∈ GF(2m) is defined by: A−1(x) · A(x) ≡ 1 mod P(x) For small fields — in practice, this often means fields with 216 or fewer elements — lookup tables which contain the precomputed inverses of all field elements are often used. Table 4.2 shows the values which are used within the S-box of AES. The table contains all inverses in GF(28) modulo P(x) = x8 + x4 + x3 + x + 1 in hexadecimal notation. A special case is the entry for the field element 0, for which an inverse does not exist. However, for the AES S-box, a substitution table is needed that is defined for every possible input value. Hence, the designers defined the S-box such that the input value 0 is mapped to the output value 0. Table 4.2 Multiplicative inverse table in GF(28) for bytes xy used within the AES S-box, with the irreducible polynomial P(x) = x8 + x4 + x3 + x + 1 Y 0 1 2 3 4 5 6 7 8 9 A B C D E F 0 00 01 8D F6 CB 52 7B D1 E8 4F 29 C0 B0 E1 E5 C7 1 74 B4 AA 4B 99 2B 60 5F 58 3F FD CC FF 40 EE B2 2 3A 6E 5A F1 55 4D A8 C9 C1 0A 98 15 30 44 A2 C2 3 2C 45 92 6C F3 39 66 42 F2 35 20 6F 77 BB 59 19 4 1D FE 37 67 2D 31 F5 69 A7 64 AB 13 54 25 E9 09 5 ED 5C 05 CA 4C 24 87 BF 18 3E 22 F0 51 EC 61 17 6 16 5E AF D3 49 A6 36 43 F4 47 91 DF 33 93 21 3B 7 79 B7 97 85 10 B5 BA 3C B6 70 D0 06 A1 FA 81 82 X 8 83 7E 7F 80 96 73 BE 56 9B 9E 95 D9 F7 02 B9 A4 9 DE 6A 32 6D D8 8A 84 72 2A 14 9F 88 F9 DC 89 9A A FB 7C 2E C3 8F B8 65 48 26 C8 12 4A CE E7 D2 62 B 0C E0 1F EF 11 75 78 71 A5 8E 76 3D BD BC 86 57 C 0B 28 2F A3 DA D4 E4 0F A9 27 53 04 1B FC AC E6 D 7A 07 AE 63 C5 DB E2 EA 94 8B C4 D5 9D F8 90 6B E B1 0D D6 EB C6 0E CF AD 08 4E D7 E3 5D 50 1E B3 F 5B 23 38 34 68 46 03 8C DD 9C 7D A0 CD 1A 41 1C Example 4.7. From Table 4.2 the inverse of x7 + x6 + x = (1100 0010)2 = (C2)hex = (xy) is given by the element in row C, column 2: (2F)hex = (0010 1111)2 = x5 + x3 + x2 + x + 1
124 4 The Advanced Encryption Standard (AES) This can be verified by multiplication: (x7 + x6 + x) · (x5 + x3 + x2 + x + 1) ≡ 1 mod P(x) Note that the table above does not contain the S-box itself, which is a bit more complex and will be described in Section 4.4.1. As an alternative to using lookup tables, one can also explicitly compute inverses. The main algorithm for computing multiplicative inverses is the extended Euclidean algorithm, which is introduced in Section 6.3.1. 4.4 Internal Structure of AES In the following, we examine the internal structure of AES. Figure 4.3 shows the block diagram of a single AES round. The 16-byte input (A0, . . . , A15) is fed byte- wise into the S-box. The 16-byte output (B0, . . . , B15) is permuted byte-wise in the ShiftRows layer and mixed by the MixColumn transformation c(x). Finally, the 128-bit subkey ki is XORed with the intermediate result. We note that AES is a byte-oriented cipher. This is in contrast to DES, which makes heavy use of bit per- mutation and can thus be considered to have a bit-oriented structure. In order to understand how the data moves through AES, we first imagine that the state A (i.e., the 128-bit data path) consists of 16 bytes A0, A1, . . . , A15, which are arranged in a four-by-four byte matrix: A0 A4 A8 A12 A1 A5 A9 A13 A2 A6 A10 A14 A3 A7 A11 A15 As we will see in the following, AES operates on elements, columns or rows of the current state matrix. Similarly, the key bytes are arranged into a matrix with four rows and four (128-bit key), six (192-bit key) or eight (256-bit key) columns. As an example, here is the state matrix of a 192-bit key: K0 K4 K8 K12 K16 K20 K1 K5 K9 K13 K17 K21 K2 K6 K10 K14 K18 K22 K3 K7 K11 K15 K19 K23 We discuss now what happens in each of the layers.
4.4 Internal Structure of AES 125Key Addition s s s s s s s s s s s s s s s s 0A 1A 2A 3A 6A5A4A 7A 8A 9A 11A 12A 13A 14A 15A10A 0B 1B 2B 3B 4B 5B 6B 7B 8B 9B 10B 11B 12B 13B 14B 15B 1C 2C 3C 4C 5C 6C 7C 8C 9C 10C 11C 12C 13C 14C 15C0C k i Byte Substitution MixColumn ShiftRows Fig. 4.3 AES round function for rounds 1, 2, . . . , nr − 1 4.4.1 Byte Substitution Layer As shown in Figure 4.3, the first layer in each round is the Byte Substitution layer. The Byte Substitution layer can be viewed as a row of 16 parallel S-boxes, each with 8 input and output bits. Note that all 16 S-boxes are identical, unlike DES where eight different S-boxes are used. In the layer, each state byte Ai is replaced, i.e., substituted, by another byte Bi: S(Ai) = Bi The S-box is the only nonlinear element of AES, i.e., it holds that ByteSub(A) + ByteSub(B) 6 = ByteSub(A + B) for two states A and B. The S-box substitution is a bijective mapping, i.e., each of the 28 = 256 possible input elements is one-to-one mapped to one output element. This allows us to uniquely reverse the S-box, which is needed for decryption. In
126 4 The Advanced Encryption Standard (AES) software implementations, the S-box is usually realized as a 256-by-8-bit lookup table with fixed entries, as given in Table 4.3. Example 4.8. Let’s assume the input byte to the S-box is Ai = (C2)hex, then the substituted value is S((C2)hex) = (25)hex On a bit level — and remember, the only thing that is ultimately of interest in en- cryption is the manipulation of bits — this substitution can be described as: S(1100 0010) = (0010 0101) Table 4.3 AES S-box: Substitution values in hexadecimal notation for input byte (xy) y 0 1 2 3 4 5 6 7 8 9 A B C D E F 0 63 7C 77 7B F2 6B 6F C5 30 01 67 2B FE D7 AB 76 1 CA 82 C9 7D FA 59 47 F0 AD D4 A2 AF 9C A4 72 C0 2 B7 FD 93 26 36 3F F7 CC 34 A5 E5 F1 71 D8 31 15 3 04 C7 23 C3 18 96 05 9A 07 12 80 E2 EB 27 B2 75 4 09 83 2C 1A 1B 6E 5A A0 52 3B D6 B3 29 E3 2F 84 5 53 D1 00 ED 20 FC B1 5B 6A CB BE 39 4A 4C 58 CF 6 D0 EF AA FB 43 4D 33 85 45 F9 02 7F 50 3C 9F A8 7 51 A3 40 8F 92 9D 38 F5 BC B6 DA 21 10 FF F3 D2 x 8 CD 0C 13 EC 5F 97 44 17 C4 A7 7E 3D 64 5D 19 73 9 60 81 4F DC 22 2A 90 88 46 EE B8 14 DE 5E 0B DB A E0 32 3A 0A 49 06 24 5C C2 D3 AC 62 91 95 E4 79 B E7 C8 37 6D 8D D5 4E A9 6C 56 F4 EA 65 7A AE 08 C BA 78 25 2E 1C A6 B4 C6 E8 DD 74 1F 4B BD 8B 8A D 70 3E B5 66 48 03 F6 0E 61 35 57 B9 86 C1 1D 9E E E1 F8 98 11 69 D9 8E 94 9B 1E 87 E9 CE 55 28 DF F 8C A1 89 0D BF E6 42 68 41 99 2D 0F B0 54 BB 16 Even though the S-box is bijective, it does not have any fixed points, i.e., there aren’t any input values Ai such that S(Ai) = Ai. Even the zero-input is not a fixed point: S(0000 0000) = (63)hex = (0110 0011). Example 4.9. Let’s assume the 16-byte input to the Byte Substitution layer is (C2,C2, . . . ,C2) in hexadecimal notation. The output state is then (25, 25, . . . , 25) Mathematical description of the S-box For readers who are interested in how the S-box entries are constructed, a more detailed description follows now. This description, however, is not necessary for a basic understanding of AES, and the
4.4 Internal Structure of AES 127 remainder of this subsection can be skipped without problem. Unlike the DES S- boxes, which are essentially random tables that fulfill certain properties, the AES S-boxes have a strong algebraic structure. An AES S-box can be viewed as a two- step mathematical transformation, as shown in Figure 4.4. Fig. 4.4 The two operations within the AES S-box which computes the function Bi = S(Ai) The first part of the substitution is a Galois field inversion, the mathematics of which were introduced in Section 4.3.2. For each input element Ai, the inverse is computed: B′ i = A−1 i , where both Ai and B′ i are considered elements in the field GF(28) with the fixed irreducible polynomial P(x) = x8 + x4 + x3 + x + 1. A lookup table with all inverses is shown in Table 4.2. Note that the inverse of the zero element does not exist. However, for AES it is defined that the zero element Ai = 0 is mapped to itself. In the second part of the substitution, each byte B′ i is multiplied by a constant bit- matrix followed by the addition of a constant 8-bit vector. The operation is described by: b0 b1 b2 b3 b4 b5 b6 b7 ≡ 1 0 0 0 1 1 1 1 1 1 0 0 0 1 1 1 1 1 1 0 0 0 1 1 1 1 1 1 0 0 0 1 1 1 1 1 1 0 0 0 0 1 1 1 1 1 0 0 0 0 1 1 1 1 1 0 0 0 0 1 1 1 1 1 b′ 0 b′ 1 b′ 2 b′ 3 b′ 4 b′ 5 b′ 6 b′ 7 + 1 1 0 0 0 1 1 0 mod 2 Note that B′ = (b′ 7, . . . , b′ 0) is the bitwise vector representation of B′ i(x) = A−1 i (x). This second step is referred to as an affine mapping. Let’s look at an example of how the S-box computations work. Example 4.10. We assume the S-box input Ai = (1100 0010)2 = (C2)hex. From Ta- ble 4.2, we can see that the inverse is: A−1 i = B′ i = (2F)hex = (0010 1111)2 We now apply the B′ i bit vector as input to the affine transformation. Note that the least significant bit b′ 0 of B′ i is at the rightmost position,
128 4 The Advanced Encryption Standard (AES) Bi = (0010 0101) = (25)hex Thus, S((C2)hex) = (25)hex, which is exactly the result that is also given in the S-box Table 4.3. If one computes both steps for all 256 possible input elements of the S-box and stores the results, one obtains Table 4.3. In most AES implementations, in particular, in virtually all software realizations of AES, the S-box outputs are not explicitly computed as shown here, but rather lookup tables like Table 4.3 are used. However, for hardware implementations, it is sometimes advantageous to realize the S-boxes as digital circuits that actually compute the inverse followed by the affine mapping. The advantage of using inversion in GF(28) as the core function of the Byte Substitution layer is that it provides a high degree of nonlinearity, which in turn provides optimum protection against some of the strongest known analytical attacks. The affine step “destroys” the algebraic structure of the Galois field, which in turn is needed to prevent attacks that would exploit the finite field inversion. 4.4.2 Diffusion Layer In AES, the Diffusion layer consists of two sublayers: the ShiftRows transformation and the MixColumn transformation. We recall that diffusion is the spreading of the influence of individual bits over the entire state. Unlike the nonlinear S-box, the diffusion layer performs a linear operation on state matrices A, B, i.e., DIFF(A) + DIFF(B) = DIFF(A + B). ShiftRows Sublayer The ShiftRows transformation cyclically shifts the second row of the state matrix by three bytes to the right, the third row by two bytes to the right and the fourth row by one byte to the right. The first row is not changed by the ShiftRows trans- formation. The purpose of the ShiftRows transformation is to increase the diffusion properties of AES. If the input of the ShiftRows sublayer is given as a state matrix B = (B0, B1, . . . , B15): B0 B4 B8 B12 B1 B5 B9 B13 B2 B6 B10 B14 B3 B7 B11 B15
4.4 Internal Structure of AES 129 The output is the new state: B0 B4 B8 B12 no shift B5 B9 B13 B1 −→ three positions right shift B10 B14 B2 B6 −→ two positions right shift B15 B3 B7 B11 −→ one position right shift (4.1) MixColumn Sublayer The MixColumn step is a linear transformation that mixes each column of the state matrix. Since every input byte influences four output bytes, the MixColumn opera- tion is the major diffusion element in AES. The combination of the ShiftRows and MixColumn sublayers makes it possible that after only three rounds, every byte of the state matrix depends on all 16 plaintext bytes. In the following, we denote the 16-byte input state by B and the 16-byte output state by C: MixColumn(B) = C where B is the state after the ShiftRows operation as given in Expression (4.1). The reader may also want to have a look again at Figure 4.3. Now, each 4-byte column of (4.1) is considered a vector and multiplied by a fixed 4 × 4 matrix. The matrix contains constant entries. Multiplication and addition of the coefficients is done in GF(28). As an example, we show how the first four output bytes are computed: C0 C1 C2 C3 = 02 03 01 01 01 02 03 01 01 01 02 03 03 01 01 02 B0 B5 B10 B15 (4.2) The second column of output bytes (C4,C5,C6,C7) is computed by multiplying the four input bytes (B4, B9, B14, B3) by the same constant matrix, and so on. Fig- ure 4.3 shows which input bytes are used in each of the four MixColumn operations. We now discuss the details of the vector–matrix multiplication, which forms the MixColumn operations. We recall that each state byte Ci and Bi is an 8-bit value representing an element from GF(28). All arithmetic involving the coefficients is done in this Galois field. For the constants in the matrix a hexadecimal notation is used: “01” refers to the GF(28) polynomial with the coefficients (0000 0001), i.e., it is the element 1 of the Galois field; “02” refers to the polynomial with the bit vector (0000 0010), i.e., to the polynomial x; and “03” refers to the polynomial with the bit vector (0000 0011), i.e., the Galois field element x + 1. The additions in the vector–matrix multiplication are GF(28) additions that are simple bitwise XORs of the respective bytes. For the multiplication of the constants, we have to realize multiplications with the constants 01, 02 and 03. These are quite efficient, and in fact, the three constants were chosen such that software implemen-
130 4 The Advanced Encryption Standard (AES) tation is easy. Multiplication by 01 is multiplication by the identity and does not involve any explicit operation. Multiplication by 02 and 03 can be done through table look-up in two 256-by-8 tables. As an alternative, multiplication by 02 can also be implemented as a multiplication by x, which is a left shift by one bit, and a modular reduction with P(x) = x8 + x4 + x3 + x + 1. Similarly, multiplication by 03, which represents the polynomial (x + 1), can be implemented by a left shift by one bit and the addition of the original value followed by a modular reduction with P(x). Example 4.11. We consider the leftmost MixColumn box in Figure 4.3 with the four input bytes: B0 = x3 + x2 ; B5 = x7 + 1 ; B10 = x5 + x4 + 1 ; B15 = x5 + x4 + x3 We show how the first output byte: C0 = 02 B0 + 03 B5 + 01 B10 + 01B15 is computed. First we precompute the multiplications 02 B0 and 03 B5: 02 B0 = x (x3 + x2) = x4 + x3 03 B5 = (x + 1) (x7 + 1) = x8 + x7 + x + 1 ≡ x7 + x4 + x3 mod P(x) C0 follows now from adding the four terms: 02 · B0 = x4+ x3 03 · B5 = x7+ x4+ x3 01 · B10 = x5+ x4+ 1 01 · B15 = x5+ x4+ x3 C0 = x7+ x3+ 1 The other output bytes can be computed the same way according to Equation (4.2), which results in C1 = x6 + x5 + x4 + x3 + x2 + x, C2 = x7 + x5 + x2 + x + 1 and C3 = x7 + x6 + x4 + x2. 4.4.3 Key Addition Layer The two inputs to the Key Addition layer are the current 16-byte state matrix and a subkey which also consists of 16 bytes (128 bits). The two inputs are combined through a bitwise XOR operation. Note that the XOR operation is equal to addi- tion in the Galois field GF(2). The subkeys are derived in the key schedule that is described below in Section 4.4.4.
4.4 Internal Structure of AES 131 4.4.4 Key Schedule The key schedule takes the original input key (of length 128, 192 or 256 bits) and derives the subkeys used in AES. Note that an XOR addition of a subkey is used both at the input and output of AES. This process is sometimes referred to as key whitening. The number of subkeys is equal to the number of rounds plus one, due to the key needed for key whitening in the first key addition layer, cf. Figure 4.2. Thus, for the key length of 128 bits, the number of rounds is nr = 10 and there are 11 subkeys, each of 128 bits. AES with a 192-bit key requires 13 subkeys of length 128 bits each, and AES with a 256-bit key has 15 subkeys. The AES subkeys are computed recursively, i.e., in order to derive subkey ki, subkey ki−1 must be known, etc. The AES key schedule is word-oriented, where one word equals 32 bits. Subkeys are stored in a key expansion array W that consists of words. There are different key schedules for the three different AES key sizes of 128, 192 and 256 bits, which are all fairly similar. We introduce the three key schedules in the following. Key Schedule for 128-Bit Key AES Each 128-bit subkey consists of 4 words. The 11 subkeys are stored in a key expan- sion array with the 11×4 = 44 elements W [0], . . . ,W [43]. The subkeys are computed as depicted in Figure 4.5. The elements K0, . . . , K15 denote the bytes of the original AES key. First, we note that the first subkey k0 is the original AES key, i.e., the key is copied into the first four elements of the key array W . The other array elements are computed as follows. As can be seen in the figure, the leftmost word of a subkey W [4i], where i = 1, . . . , 10, is computed as: W [4i] = W [4(i − 1)] ⊕ g(W [4i − 1]) Here g is a nonlinear function with a four-byte input and output. The remaining three words of a subkey are computed recursively as: W [4i + j] = W [4i + j − 1] ⊕W [4(i − 1) + j] where i = 1, . . . , 10 and j = 1, 2, 3. The function g rotates its four input bytes, per- forms a byte-wise S-box substitution, and adds a round coefficient RC to it. The round coefficient is an element of the Galois field GF(28), i.e., an 8-bit value. It is only added to the leftmost byte in the function g.
132 4 The Advanced Encryption Standard (AES)0 K 3 K 4 K 7 K 8 K 11 K 12 K 15 K gfunction of round 32 8 RC[i] 3V2V1V0V SSSS 8 32 2V 0V1V 3V W[7]W[6]W[5]W[4] W[3]W[2]W[1]W[0] W[43]W[42]W[41]W[40] W[39]W[38]W[37]W[36] 32 32 32 32 g round key 10 round key 9 round key 1 round key 0 g .... .... .... i Fig. 4.5 AES key schedule for 128-bit key size The round coefficients vary from round to round according to the following rule: RC[1] = x0 = (0000 0001)2 RC[2] = x1 = (0000 0010)2 RC[3] = x2 = (0000 0100)2 ... RC[10] = x9 = (0011 0110)2 The function g has two purposes. First, it adds nonlinearity to the key sched- ule. Second, it removes symmetry in AES. Both properties are necessary to thwart certain block cipher attacks.
4.4 Internal Structure of AES 133 Key Schedule for 192-Bit Key AES AES with a 192-bit key has 12 rounds and, thus, 13 subkeys of 128 bits each. The subkeys require 13 × 4 = 52 words, which are stored in the array elements W [0], . . . ,W [51]. The computation of the array elements is quite similar to the 128- bit key case and is shown in Figure 4.6. There are eight rounds of the key schedule.gfunction of round 32 8 RC[i] 3V2V1V0V SSSS 8 32 2V 0V1V 3V W[51]W[50]W[49]W[48] W[47]W[46]W[45]W[44]W[43]W[42] W[11]W[10]W[9]W[8]W[7]W[6] W[5]W[4]W[3]W[2]W[1]W[0] g .... .... .... .... .... g 323232 32 32 32 0 K 3 K 4 K 7 K 8 K 11 K 12 K 15 K 16 K 19 K 20 K 23 K i Fig. 4.6 AES key schedule for 192-bit key size (Note that these key schedule rounds do not correspond to the 12 AES rounds!) Each iteration computes six new words of the subkey array W . The subkey for the first AES round is formed by the array elements (W [0],W [1],W [2],W [3]), the second subkey by the elements (W [4],W [5],W [6],W [7]) and so on. Eight round coefficients RC[i] are needed within the function g. They are computed as in the 128-bit case and range from RC[1], . . . , RC[8].
134 4 The Advanced Encryption Standard (AES) Key Schedule for 256-Bit Key AES AES with a 256-bit key needs 15 subkeys. The subkeys are stored in the 15 × 4 = 60 words W [0], . . . ,W [59]. The computation of the array elements is quite similar to the 128-bit key case and is shown in Figure 4.7. The key schedule has seven0 K 3 K 4 K 7 K 8 K 11 K 12 K 15 K 16 K 19 K 20 K 23 K 24 K 27 K 28 K 31 K W[0] W[56] W[57] W[58] W[59] W[8] W[9] W[10] W[11] W[12] W[13] W[14] W[15] g W[48] W[49] W[50] W[51] W[52] W[53] W[54] W[55] g h 32 V0 V1 V2 V3 S S S S 32 −functionh V3V1 V0V2 32 8 S S S S V0 V1 V2 V3 32 RC[i] 8 function of roundg i 32323232 32 32 32 32 .... .... .... .... .... .... .... W[7]W[6]W[5]W[4]W[3]W[2]W[1] Fig. 4.7 AES key schedule for 256-bit key size rounds, where each round computes eight words for the subkeys. (Again, note that these key schedule rounds do not correspond to the 14 AES rounds.) The subkey for the first AES round is formed by the array elements (W [0],W [1],W [2],W [3]),
4.5 Decryption 135 the second subkey by the elements (W [4],W [5],W [6],W [7]) and so on. Seven round coefficients RC[1], . . . , RC[7] within the function g are needed within and they are computed as in the 128-bit case. This key schedule also has a function h with a 4-byte input and output. The function applies the S-box to all four input bytes. In general, when implementing any of the three key schedules, two different approaches exist: 1. Precomputation All subkeys are expanded first into the array W . The encryp- tion (decryption) of a plaintext (ciphertext) is executed afterward. This approach is often taken in PC and server implementations of AES, where large pieces of data are encrypted under one key. Please note that this approach requires (nr + 1) · 16 bytes of memory, e.g., 11 · 16 = 176 bytes if the key size is 128 bits. There are applications with very limited memory, such as small IoT-like devices, where this precomputation can be difficult. 2. On-the-fly A new subkey is derived for every new AES round during the en- cryption (decryption) of a plaintext (ciphertext). Please note that when decrypting ciphertexts, the last subkey is XORed first with the ciphertext. Therefore, it is re- quired to recursively derive all subkeys first and then begin the decryption of the ciphertext and the on-the-fly generation of subkeys. As a result of this overhead, the decryption of a ciphertext is always slightly slower than the encryption of a plaintext when on-the-fly generation of subkeys is used. 4.5 Decryption Because AES is not based on a Feistel network, all layers must actually be inverted for the decryption, i.e., the Byte Substitution layer becomes the Inv Byte Substitu- tion layer, the ShiftRows layer becomes the Inv ShiftRows layer, and the MixCol- umn layer becomes the Inv MixColumn layer. However, as we will see, it turns out that the inverse layer operations are fairly similar to the layer operations used for encryption. In addition, the order of the subkeys is reversed, i.e., we need a reversed key schedule. A block diagram of the decryption function is shown in Figure 4.8. Since the last encryption round does not perform the MixColumn operation, the first decryption round also does not contain the corresponding inverse layer. All other decryption rounds, however, contain all AES layers. In the following, we dis- cuss the inverse layers of the general AES decryption round (Figure 4.9). Since the XOR operation is its own inverse, the key addition layer in the decryption mode is the same as in the encryption mode: It consists of a row of plain XOR gates.
136 4 The Advanced Encryption Standard (AES)Plaintext nTransform −1 1Transform 0Transform rn inverse of round nr inverse of round −1nr Inv Byte Substitution Inv MixColumn Layer Inv ShiftRows Layer Key Addition Layer Key Addition Layer inverse of round 1 Inv Byte Substitution Inv MixColumn Layer Inv ShiftRows Layer Key Addition Layer Key Addition Layer Inv ShiftRows Layer Inv Byte Substitution Ciphertext y AES ( ) k0 k1 k r −1n knr Transform −1 kKey x= y r Fig. 4.8 AES decryption block diagram
4.5 Decryption 137InvSubBytes 1C 2C 3C 4C 5C 6C 7C 8C 9C 10C 11C 12C 13C 14C 15C0C k i 10B 7B 14B1B4B 11B 8B 5B 15B 12B 9B 6B 3B2B0B 13B 0A 1A 3A 5A 6A 8A 9A 10A 11A 12A 13A 14A 15A2A 4A 7A s−1 s−1 s−1 s−1 s−1 s−1 s−1 s−1 s−1 s−1 s−1 s−1 s−1 s−1 s−1 s−1 InvMixColumn InvShiftRows Key Addition Fig. 4.9 AES decryption round function Inverse MixColumn Sublayer After the addition of the subkey, the inverse MixColumn step is applied to the state (again, the exception is the first decryption round). In order to reverse the MixCol- umn operation, the inverse of its matrix must be used. The input is a 4-byte column of the state C, which is multiplied by the inverse 4 × 4 matrix. The matrix contains constant entries. Multiplication and addition of the coefficients is done in GF(28). Below is the vector-matrix multiplication for the leftmost MixColumn box in Fig- ure 4.9: B0 B1 B2 B3 = 0E 0B 0D 09 09 0E 0B 0D 0D 09 0E 0B 0B 0D 09 0E C0 C1 C2 C3
138 4 The Advanced Encryption Standard (AES) The second column of output bytes (B4, B5, B6, B7) is computed by multiplying the four input bytes (C4,C5,C6,C7) by the same constant matrix, and so on. Each value Bi and Ci is an element from GF(28). Also, the constants are elements from GF(28). The notation for the constants is hexadecimal and is the same as was used for the MixColumn layer, for example: 0B = (0B)hex = (0000 1011)2 = x3 + x + 1 Additions in the vector–matrix multiplication are bitwise XORs. Inverse ShiftRows Sublayer In order to reverse the ShiftRows operation of the encryption algorithm, we must shift the rows of the state matrix in the opposite direction. The first row is not changed by the inverse ShiftRows transformation. If the input of the ShiftRows sublayer is given as a state matrix B = (B0, B1, . . . , B15): B0 B4 B8 B12 B1 B5 B9 B13 B2 B6 B10 B14 B3 B7 B11 B15 then the inverse ShiftRows sublayer yields the output: B0 B4 B8 B12 no shift B13 B1 B5 B9 ←− three positions left shift B10 B14 B2 B6 ←− two positions left shift B7 B11 B15 B3 ←− one position left shift Inverse Byte Substitution Layer The inverse S-box must used when decrypting a ciphertext. Since the AES S-box is a bijective, i.e., a one-to-one mapping, it is possible to construct an inverse S-box such that: Ai = S−1(Bi) = S−1(S(Ai)) where Ai and Bi are elements of the state matrix. The entries of the inverse S-box are given in Table 4.4. For readers who are interested in the details of how the entries of the inverse S-box are constructed, we provide a derivation below. However, for a functional understanding of AES, the remainder of this section can be skipped. In order to reverse the S-box substitution, we first have to compute the inverse of the affine transformation, cf. Figure 4.4. For this, each input byte Bi is considered an element of GF(28). The inverse affine transformation on each byte Bi is given by
4.5 Decryption 139 Table 4.4 Inverse AES S-box: Substitution values in hexadecimal notation for input byte (xy) y 0 1 2 3 4 5 6 7 8 9 A B C D E F 0 52 09 6A D5 30 36 A5 38 BF 40 A3 9E 81 F3 D7 FB 1 7C E3 39 82 9B 2F FF 87 34 8E 43 44 C4 DE E9 CB 2 54 7B 94 32 A6 C2 23 3D EE 4C 95 0B 42 FA C3 4E 3 08 2E A1 66 28 D9 24 B2 76 5B A2 49 6D 8B D1 25 4 72 F8 F6 64 86 68 98 16 D4 A4 5C CC 5D 65 B6 92 5 6C 70 48 50 FD ED B9 DA 5E 15 46 57 A7 8D 9D 84 6 90 D8 AB 00 8C BC D3 0A F7 E4 58 05 B8 B3 45 06 7 D0 2C 1E 8F CA 3F 0F 02 C1 AF BD 03 01 13 8A 6B x 8 3A 91 11 41 4F 67 DC EA 97 F2 CF CE F0 B4 E6 73 9 96 AC 74 22 E7 AD 35 85 E2 F9 37 E8 1C 75 DF 6E A 47 F1 1A 71 1D 29 C5 89 6F B7 62 0E AA 18 BE 1B B FC 56 3E 4B C6 D2 79 20 9A DB C0 FE 78 CD 5A F4 C 1F DD A8 33 88 07 C7 31 B1 12 10 59 27 80 EC 5F D 60 51 7F A9 19 B5 4A 0D 2D E5 7A 9F 93 C9 9C EF E A0 E0 3B 4D AE 2A F5 B0 C8 EB BB 3C 83 53 99 61 F 17 2B 04 7E BA 77 D6 26 E1 69 14 63 55 21 0C 7D Equation (4.3), where (b7, . . . , b0) is the bitwise vector representation of Bi(x), and (b′ 7, . . . , b′ 0) the result after the inverse affine transformation: b′ 0 b′ 1 b′ 2 b′ 3 b′ 4 b′ 5 b′ 6 b′ 7 ≡ 0 0 1 0 0 1 0 1 1 0 0 1 0 0 1 0 0 1 0 0 1 0 0 1 1 0 1 0 0 1 0 0 0 1 0 1 0 0 1 0 0 0 1 0 1 0 0 1 1 0 0 1 0 1 0 0 0 1 0 0 1 0 1 0 b0 b1 b2 b3 b4 b5 b6 b7 + 1 0 1 0 0 0 0 0 mod 2 (4.3) In the second step of the inverse S-box operation, the Galois field inverse has to be reversed. For this, note that Ai = (A−1 i )−1. This means that the inverse operation is reversed by computing the inverse again. In our notation we thus have to compute Ai = (B′ i)−1 ∈ GF(28) with the fixed reduction polynomial P(x) = x8 + x4 + x3 + x + 1. Again, the zero ele- ment is mapped to itself. The vector Ai = (a7, . . . , a0) (representing the field element a7x7 + · · · + a1x + a0) is the result of the substitution: Ai = S−1(Bi) Decryption Key Schedule Since decryption round one needs the last subkey, the second decryption round needs the second-to-last subkey and so on, we need the nr + 1 subkeys in reversed
140 4 The Advanced Encryption Standard (AES) order, as shown in Figure 4.8. In practice, this is mainly achieved by computing the entire key schedule first and storing all 11, 13 or 15 subkeys, depending on the number of rounds AES is using (which in turn depends on the three key lengths sup- ported by AES). This precomputation usually adds a small latency to the decryption operation relative to encryption. 4.6 Implementation in Software and Hardware Below, we briefly comment on the efficiency of the AES cipher with respect to software and hardware implementation. Software Unlike DES, AES has a much more software-friendly design. A straightforward implementation of AES, which directly follows the data path-oriented description provided in this chapter, is well suited for 8-bit processors, which are sometimes used in simple smart cards or low-cost IoT devices. For 32- and 64-bit CPUs, which are common in today’s smartphones, PCs and servers, table-based implementations are widely used. The core idea is to merge all round functions (except the rather triv- ial key addition) into one table look-up. This results in four tables, each of which consists of 256 entries, where each entry is 32 bits wide. These tables are named T-Boxes. Interestingly, this table method was already described by the Rijndael de- signers in the original AES documentation submitted during the NIST competi- tion. Four table accesses yield the 32 output bits of one round. Hence, one round can be computed with 16 table look-ups. On a 1.2-GHz Intel processor, a through- put of 50 MByte/s (or 400 Mbit/s) is possible. On many processors from Intel and AMD, special instructions, called AES-NI (Advanced Encryption Standard New In- structions), are available, which accelerate the computation of AES. AES-NI allows throughput in the range of 2 GByte/s (or 16 Gbit/s). Hardware Compared to DES, AES requires more hardware resources for an implementation. However, due to the high integration density of modern integrated circuits, hardware circuits for AES are still quite small compared to many other functions commonly implemented in modern integrated circuits. Due to its wide data path of 128 bits, which can be implemented in parallel, AES hardware realizations are possible with very high throughputs in modern ASIC or FPGA (field programmable gate array — these are programmable hardware devices) technology. Implementations of AES engines that use a single round that is iterated 10, 12 or 14 times (depending on the key length) can exceed throughputs of 10 Gbit/sec. By implementing several such
4.7 Discussion and Further Reading 141 rounds in parallel and pipelining them, the speed can be further increased. It can be said that symmetric encryption with today’s ciphers is extremely fast, not only com- pared to asymmetric cryptosystems but also compared to other algorithms needed in modern communication systems, such as data compression or signal processing schemes. 4.7 Discussion and Further Reading AES Algorithm and Security A detailed description of the design principles of AES can be found in the book [87], which was written by the two Rijndael inventors. More than two decades after it was standardized by NIST, AES is today extremely widely used in practice. As mentioned at the very beginning of this chapter, AES is part of the web security protocol TLS, the internet security standard IPsec, the Wi-Fi encryption standard IEEE 802.11i, the secure shell network protocol SSH, the instant messengers WhatsApp and Signal, many hard disk encryption products and numerous security products around the world. Besides its unique position as probably the most-used cipher worldwide, AES triggered very influential ideas on how to design secure block ciphers. A prominent example is the wide trail strategy [86], which demonstrates and clarifies the role of the linear layer of a block cipher with respect to its resistance against many statistical attacks. It is probably fair to say that the majority of today’s successful block ciphers have borrowed ideas from AES. Currently no analytical attack against AES is known which has a complexity less than a brute-force attack. An elegant algebraic description was found in [192], which in turn triggered speculations that this could lead to attacks. Subsequent research showed that an attack is, in fact, not feasible. By now, the common assumption is that the approach will not threaten AES. A good summary on algebraic attacks can be found in [72]. In addition, there have been proposals for many other attacks, including the square attack, impossible differential attack or related-key attack. The best attacks on full-round AES are probably the biclique attacks [58]. They can be seen as attacks that improve the complexity of brute-force attacks by not having to iterate over the entire AES encryption for every possible key but rather iterate over parts only. Biclique attacks marginally improve the complexity of a brute-force attack by a factor of about four, e.g., attacking AES-128 takes about 2126 steps. None of the proposed attacks come even close to threatening full-size AES in practical settings.. Galois Fields The standard reference for the mathematics of finite fields is [174]. A very accessible but brief introduction is also given in [46]. The International Work- shop on the Arithmetic of Finite Fields (WAIFI) is concerned with both the applica- tions and the theory of Galois fields [250]. AES Implementations As mentioned in Section 4.6, many software implemen- tations use special lookup tables (T-Boxes) for realizing AES. An early detailed
142 4 The Advanced Encryption Standard (AES) description of the construction of T-Boxes can be found in [85, Section 5]. A de- scription of a high-speed software implementation on 32-bit and 64-bit CPUs is given in [183, 182]. The bit-slicing technique, which was developed in the context of DES, is also applicable to AES and can lead to very fast code, as shown in [184]. A strong indication of the importance of AES was the introduction of special AES instructions for CPUs from AMD and Intel. They are extensions to the x86 in- struction set architecture, referred to as Advanced Encryption Standard New Instruc- tions (short AES-NI), originally proposed by Intel in 2008 [129]. The instructions allow these machines to compute the AES round operations particularly quickly. There is a wealth of literature dealing with hardware implementation of AES. A good introduction to the area of AES hardware architectures is given in [165, Chap- ter 10]. As an example of the variety of AES implementations, reference [127] de- scribes a very small FPGA implementation with 2.2 Mbit/s and a very fast pipelined FPGA implementation with 25 Gbit/s. It is also possible to use the DSP blocks (i.e., fast arithmetic units) available on modern FPGAs for AES, which can also yield throughputs beyond 50 Gbit/s [99]. The basic idea in all high-speed architectures is to process several plaintext blocks in parallel by means of pipelining. At the other end of the performance spectrum are lightweight architectures that are optimized for applications such as small IoT devices. The basic idea here is to serialize the data path, i.e., one round is processed in several time steps. Good references are [113, 71]. 4.8 Lessons Learned AES is a modern block cipher that supports three key lengths of 128, 192 and 256 bits. It provides excellent long-term security against brute-force attacks. AES has been studied intensively since the late 1990s and no attacks have been found that are more than marginally better than brute-force. AES is not based on Feistel networks. The AES rounds make heavily use of Galois field arithmetic. AES has an excellent performance in software and hardware. AES is part of numerous open standards, such as IPsec or TLS, in addition to being the mandatory encryption algorithm for U.S. government applications. It seems likely that the cipher will continue to be the dominant symmetric encryp- tion algorithm for many years to come. AES is efficient in software and hardware.
4.8 Problems 143 Problems 4.1. The AES cipher was standardized by NIST as a U.S. standard in 2001 and soon adopted by numerous international standards. It is nowadays the dominant symmet- ric cipher in use (at least in the Western world). Answer the following questions: 1. The evolution of AES differs from that of DES. Briefly describe the differences of the AES’s history in comparison to that of DES. 2. Outline the fundamental events of the development process. 3. What is the name of the algorithm that is known as AES? 4. Who developed this algorithm, and from which country are its inventors? 5. What block sizes and key lengths are supported by the cipher? 4.2. Within the AES algorithm, some computations are done in Galois fields. With the following problems, we practice some basic computations. Compute the multiplication and addition table for the prime field GF(7). A mul- tiplication table is a square table (here: 7 × 7) that has as its rows and columns all field elements. Each of its entries is the product of the field element at the corre- sponding row and column. Note that the table is symmetric along the diagonal. An addition table is completely analogous but contains the sums of field elements as entries. 4.3. Generate the multiplication table for the extension field GF(23) for the case that the irreducible polynomial is P(x) = x3 + x + 1. The multiplication table, in this case, is an 8 × 8 table. (Remark: You can do this manually or write a program for it.) 4.4. In this problem, we look at addition in extension fields. Compute A(x) + B(x) mod P(x) in GF(24) using the irreducible polynomial P(x) = x4 + x + 1. What would happen if we would change the reduction polynomial? 1. A(x) = x2 + 1, B(x) = x3 + x2 + 1. 2. A(x) = x2 + 1, B(x) = x + 1. 4.5. We now look at multiplication in the field GF(24): Compute A(x) · B(x) mod P(x) in GF(24) using the irreducible polynomial P(x) = x4 + x + 1. What is the influence of the choice of the reduction polynomial on the computation in this case? 1. A(x) = x2 + 1, B(x) = x3 + x2 + 1. 2. A(x) = x2 + 1, B(x) = x + 1. 4.6. Compute in GF(28): (x4 + x + 1)/(x7 + x6 + x3 + x2) where the irreducible polynomial is the one used by AES, namely P(x) = x8 + x4 + x3 + x + 1. Note that Table 4.2 contains a list of all multiplicative inverses for this field.
144 4 The Advanced Encryption Standard (AES) 4.7. We consider the finite field GF(24) with P(x) = x4 + x + 1 being the irreducible polynomial. Find the inverses of A(x) = x and B(x) = x2 + x. You can find the in- verses either by trial and error, i.e., brute-force search, or by applying the Euclidean algorithm for polynomials. (However, the Euclidean algorithm is only sketched briefly in this chapter.) Verify your answer by multiplying the inverses you deter- mined by A and B, respectively. 4.8. Find all irreducible polynomials 1. of degree 3 over GF(2), 2. of degree 4 over GF(2). The best approach for doing this is to consider all polynomials of lower degree and check whether you can split them into factors (other than 1 and the polynomial itself). Note that we only consider monic irreducible polynomials, i.e., polynomials with the highest coefficient equal to one. 4.9. We consider the first part of the ByteSub operation, i.e., the Galois field inver- sion. 1. Using Table 4.2, what is the inverse of the bytes 29, F3 and 01, where each byte is given in hexadecimal notation? 2. Verify your answer by performing a GF(28) multiplication with your answer and the input byte. Note that you have to represent each byte first as a polynomial in GF(28). The MSB of each byte represents the x7 coefficient. 4.10. Your task is to compute the S-box, i.e., the ByteSub, values for the input bytes 29, F3 and 01, where each byte is given in hexadecimal notation. 1. First, look up the inverses using Table 4.2 to obtain values B′. Now, perform the affine mapping by computing the matrix–vector multiplication and addition. 2. Verify your result using the S-box in Table 4.3. 3. What is the value of S(0)? 4.11. We consider AES with a 128-bit key. What is the output of the first round of AES if the input of the first Byte Substitution Layer consists of 128 ones, and the first subkey (i.e., k1) also consists of 128 ones? It is best if you write the output state as a 4 × 4 array, as shown in Section 4.4. 4.12. In the following, we check the diffusion property of AES after a single round. Let X = (X0, X1, X2, X3) = (0x01000000, 0x00000000, 0x00000000, 0x00000000) be the four input words (32 bits each) to AES with a 128-bit key. Note that all input bits except one have the value zero. The subkeys for the first round are denoted by W [0], . . . ,W [7] with 32 bits each, and are given by: W [0] = (0x2B7E1516), W [1] = (0x28AED2A6), W [2] = (0xABF71588), W [3] = (0x09CF4F3C), W [4] = (0xA0FAFE17), W [5] = (0x88542CB1), W [6] = (0x23A33939), W [7] = (0x2A6C7605).
4.8 Problems 145 Perform the following computations. You can do them manually or write a com- puter program or use an existing one. 1. Compute the output of the first round of AES for the given input X and subkeys W [0], . . . ,W [7]. Also provide all intermediate steps, i.e., the outputs after the lay- ers SubBytes, ShiftRows and MixColumns. It is recommended that you provide all results in a 4 × 4 state array. 2. Compute the output of the first round of AES for the case that all input bits are zero. 3. How many output bits have changed? Note that we only consider a single round — after every further round, more output bits will be affected. This behavior is referred to as the avalanche effect. 4.13. We are in the 9th round of an AES-128 decryption. Your task is to compute the final round and the corresponding plaintext. The AES-128 key is given by: k0 = (0x00112233, 0x44556677, 0x88990011, 0x22334455) = 00 44 88 22 11 55 99 33 22 66 00 44 33 77 11 55 The state after the inverse byte substitution in Round 9 is 77 e9 45 58 c4 c f f a c3 b8 66 87 95 05 23 31 20 For your answer, write the output state also as a 4 × 4 array. 4.14. For the following, we assume AES with a 192-bit key. Furthermore, we as- sume we have access to special-purpose integrated circuits that can check 3 · 107 AES keys per second. 1. If we use 100,000 such ICs in parallel, how long does an average key search take? Compare this period of time with the age of the universe (appr. 1010 years). 2. Assuming Moore’s law will still be valid for the next few years, how many years do we have to wait until we can build a key-search machine to perform an average key search of AES-192 in 24 hours? Again, assume that we use 100,000 ICs in parallel. 4.15. The MixColumn operation has a somewhat surprising property: If the four input bytes all have the same value, the four output bytes are also all the same, and additionally have the same value as the input byte. (We note that this property does not lead to a security weakness of AES.) 1. Show that the property holds.
146 4 The Advanced Encryption Standard (AES) 2. Compute the probability that four random input bytes take the same value. 3. Argue why this situation (i.e., all four input bytes of the MixColumn operation are identical) occurs infrequently in practice if AES is used with a random key. 4.16. Show that the two constant matrices for the MixColumn and Inv MixColumn are each other’s inverses. 4.17. In this problem we look at computations in the MixColumn layer on the bit level. As we saw in this chapter, the MixColumn transformation consists of a matrix–vector multiplication in the finite field GF(28) with the irreducible polyno- mial P(x) = x8 + x4 + x3 + x + 1. Let b = (b7x7 + . . . + b0) be one of the (four) input bytes to the vector–matrix multiplication. Within the matrix multiplication, the byte b is multiplied with the constants 01, 02 and 03. We are interested in the bit-level equations for computing those three constant multiplications. We denote the result by d = (d7x7 + . . . + d0). Your task is to express the bits di in terms of the input bits bi. 1. Equations for computing the 8 bits of d = 01 · b. 2. Equations for computing the 8 bits of d = 02 · b. 3. Equations for computing the 8 bits of d = 03 · b. Note that in the AES specifications “01” represents the polynomial 1, “02” repre- sents the polynomial x, and “03” represents x + 1. 4.18. We now look at the gate (or bit) complexity of the MixColumn function, using the results from problem 4.17. We recall from the discussion of stream ciphers that a 2-input XOR gate performs a GF(2) addition. 1. How many 2-input XOR gates are required to perform one constant multiplica- tion by 01, 02 and 03, respectively, in GF(28)? 2. What is the overall gate complexity of one matrix–vector multiplication? (The gate complexity is important when implementing AES in hardware.) 3. What is the overall gate complexity of a hardware implementation of the entire Diffusion layer? Note that permutations require no gates. 4.19. Derive the bit representation for the following round constants within the key schedule: RC[8], RC[9].
Chapter 5 More About Block Ciphers A block cipher is much more than just an encryption algorithm. It can be used as a versatile building block with which a diverse set of cryptographic mechanisms can be realized. For instance, we can use them for building different types of block- based encryption schemes, and we can even use block ciphers for realizing stream ciphers. The different ways of encryption are called modes of operation and are discussed in this chapter. Block ciphers can also be used for constructing hash func- tions, message authentication codes, which are also knowns as MACs, or key estab- lishment protocols, all of which will be described in later chapters. There are also other uses for block ciphers, e.g., as pseudorandom number generators. In addition to modes of operation, this chapter also discusses two useful techniques for increas- ing the security of block ciphers, namely key whitening and multiple encryption. In this chapter you will learn: Important modes of operation for block ciphers in practice Security pitfalls when using modes of operation The principles of key whitening Why double encryption is not a good idea due to the meet-in-the-middle attack Triple encryption 147 C. Paar et al., Understanding Cryptography, https://doi.org/10.1007/978-3-662-69007-9_5 © The Editor(s) (if applicable) and The Author(s), under exclusive license to Springer-Verlag GmbH, DE, part of Springer Nature 2024
148 5 More About Block Ciphers 5.1 Modes of Operation for Encryption and Authentication In this section, we will introduce several approaches to using block ciphers for en- crypting data, referred to as modes of operation. We will also briefly discuss modes of operation that are used for message authentication. However, the corresponding modes will be described in detail in Section 13.3. Modes of Operation for Encryption In the previous chapters we introduced how AES, DES, 3DES and PRESENT encrypt a block of data. Of course, in practice one wants typically to encrypt more than one single 8-byte or 16-byte block of plaintext, e.g., when encrypting an email or a PDF file. There are several ways of encrypting long plaintexts with a block cipher, including the following modes of operation, which will be introduced in this chapter: Electronic Code Book mode (ECB), Cipher Block Chaining mode (CBC), Cipher Feedback mode (CFB), Output Feedback mode (OFB), Counter mode (CTR), XEX Tweakable Block Cipher with Ciphertext Stealing (XTS). All modes provide confidentiality for a message sent from Alice to Bob. We note that the CFB and OFB modes use the block cipher as a building block for a stream cipher. We will also introduce XTS, a mode that has been standardized for encryption of data in storage devices such has hard disks. For completeness, we note that there are modes for other purposes too, e.g., the key wrapping modes KW, KWP and TKW, which are used when cryptographic keys need to be encrypted. The ECB and CBC modes require that the length of the plaintext is an exact multiple of the block size of the cipher used, e.g., a multiple of 16 bytes in the case of AES. If the plaintext does not have this length, it must be padded. There are several ways of doing this padding in practice. One possible padding method is to append a single “1” bit to the plaintext and then to append as many “0” bits as necessary to reach a multiple of the block length. Should the plaintext be an exact multiple of the block length, an extra block consisting only of padding bits is appended. Modes of Operation for Authentication In practice, we often not only want to keep data confidential but Bob also wants to know whether the message is really coming from Alice. This is called message authentication and the following modes of operation provide authentication (CBC-MAC, CMAC), or both encryption and authentication (CCM, GCM): CBC-MAC, Cipher-based MAC (CMAC), Cipher Block Chaining-Message Authentication Code (CCM), Galois Counter mode (GCM). These modes enable the receiving party, Bob, to determine whether the message was indeed created by a party who is in possession of the shared secret key. Moreover,
5.1 Modes of Operation for Encryption and Authentication 149 authentication also allows Bob to detect whether the ciphertext was altered during transmission. These modes are the topic of Section 13.3. 5.1.1 Electronic Codebook Mode (ECB) The Electronic Code Book (ECB) mode is the most straightforward way of encrypt- ing a message with a block cipher. In the following, let ek(xi) denote the encryption of plaintext block xi with key k using some arbitrary block cipher. Let e−1 k (yi) de- note the decryption of ciphertext block yi with key k. Let us assume that the block cipher encrypts (decrypts) blocks with a size of b bits. Messages which exceed b bits are partitioned into b-bit blocks. If the length of the message is not a multiple of b bits, it must be padded to a multiple of b bits prior to encryption. As shown in Figure 5.1, in ECB mode each block is encrypted separately. The block cipher can, for instance, be AES or 3DES. Fig. 5.1 Encryption and decryption in ECB mode Encryption and decryption in the ECB mode is formally described as follows. Definition 5.1.1 Electronic Codebook Mode (ECB) Let e() be a block cipher of block size b, and let xi and yi be bit strings of length b. Encryption: yi = ek(xi), i ≥ 1 Decryption: xi = e−1 k (yi), i ≥ 1 It is straightforward to verify the correctness of the ECB mode: e−1 k (yi) = e−1 k (ek(xi)) = xi The ECB mode has advantages. Block synchronization between the encryption and decryption parties Alice and Bob is not necessary, i.e., if the receiver does not receive all encrypted blocks due to transmission problems, it is still possible to de- crypt the received blocks. Similarly, bit errors, e.g., caused by noisy transmission
150 5 More About Block Ciphers lines, only affect the corresponding block but not succeeding blocks. Also, block ci- phers operating in ECB mode can be parallelized, e.g., one encryption unit encrypts (or decrypts) block 1, the next one block 2, and so on. This is an advantage for high-speed implementations. We note that other modes such as CFB do not allow parallelization. However, as is often the case in cryptography, there are some unexpected weak- nesses associated with the ECB mode, which we will discuss in the following. The main problem of the ECB mode is that it encrypts highly deterministically. This means that identical plaintext blocks result in identical ciphertext blocks, as long as the key does not change. The ECB mode can be viewed as a gigantic code book — hence the mode’s name — which maps every input to a certain output. Of course, if the key is changed the entire code book changes but as long as the key is static the book is fixed. This has several undesirable consequences. First, an attacker rec- ognizes whether the same message has been sent twice simply by looking at the ciphertext. Deducing information from the ciphertext in this way is called traffic analysis. For instance, if there is a fixed header that always precedes a message, the header always results in the same ciphertext. From this, an attacker can, for in- stance, learn when a message with this specific header has been sent. Second, plain- text blocks are encrypted independently of previous blocks. If an attacker reorders the ciphertext blocks, this might result in valid plaintext and the reordering might not be detected. We demonstrate two simple attacks which exploit these weaknesses of the ECB mode. The ECB mode is susceptible to a substitution attack because once a particular plaintext to ciphertext block mapping xi → yi is known, a sequence of ciphertext blocks can easily be manipulated. We demonstrate how a substitution attack could work in the real world. Imagine the following example of an electronic wire transfer betweens banks. Example 5.1. Substitution attack against electronic bank transfer Let’s assume a simple protocol for wire transfers between banks (Figure 5.2). There4 51 2 3Block # Amount $ Receiving Account # Receiving Bank B Sending Account # Sending Bank A Fig. 5.2 Simple wire-transfer protocol that can be exploited by a substitution attack against ECB encryption are five fields that specify a transfer: the sending bank’s ID and account number, the receiving bank’s ID and account number and the amount. We assume now (and this is a major simplification) that each of the fields has exactly the size of the block cipher width, e.g., 16 bytes in the case of AES. Furthermore, the encryption key
5.1 Modes of Operation for Encryption and Authentication 151 between the two banks does not change too frequently. Due to the nature of the ECB mode, an attacker can exploit the deterministic nature of this mode of operation by simple substitution of the blocks. The attack details are as follows: 1. The attacker, Oscar, opens one account at bank A and one at bank B. 2. Oscar taps the encrypted line of the banking communication network. 3. He sends $1.00 transfers from his account at bank A to his account at bank B repeatedly. He observes the ciphertexts going through the communication net- work. Even though he cannot decipher the random-looking ciphertext blocks, he can check for ciphertext blocks that repeat. After a while he can recognize the five blocks of his own transfer. He now stores blocks 1, 3 and 4 of these transfers. These are the encrypted versions of the ID numbers of both banks as well as the encrypted version of his account at bank B. 4. Recall that the two banks do not change the key too frequently. This means that the same key is used for many transfers between A and B. By comparing blocks 1 and 3 of all subsequent messages with the ones he has stored, Oscar recognizes all transfers that are made from some account at bank A to some account at bank B. He now simply replaces block 4 — which contains the receiving account number — with the block 4 that he stored before. This block contains Oscar’s account number in encrypted form. As a consequence, all transfers from some account of bank A to some account of bank B are redirected to go into Oscar’s account! Note that bank B now has no means of detecting that block 4 has been replaced in some of the transfers it receives. 5. OSCAR withdraws money from bank B quickly and flies to a country that has a relaxed attitude about the extradition of white-collar criminals. What’s interesting about this attack is that it works completely without breaking the block cipher itself. So even if we were to use AES with a 256-bit key and if we were to encrypt each block, say, 1000 times, this would not prevent the attack. Note that this attack only works if the key between bank A and B is not changed too frequently. This is another reason why key freshness is a good idea. It should be stressed that the confidentiality provided by the block cipher is not directly broken: Messages that are unknown to Oscar still remain confidential. He simply replaces parts of the ciphertext with parts of some other (previous) cipher- texts. This is called a violation of the integrity of the message. There are techniques available for preserving the message integrity, namely message authentication codes (MACs) and digital signatures. Both are widely used in practice to prevent such an attack and are introduced in Chapters 13 and 10, respectively. There are also modes of operation with built-in message authentication, referred to as authenticated en- cryption. Sections 13.3.3 and 13.3.4 describe two authenticated encryption modes. We now look at another problematic application of the ECB mode. Example 5.2. Encryption of bitmaps in ECB mode Figure 5.3 shows the original version of a bitmap image on the top with the ECB- encrypted version underneath. It is immediately obvious that there is major infor- mation leakage: The text in the graphic is still readable from the encrypted picture
152 5 More About Block Ciphers even though we used AES with a 256-bit key for encryption. The reason for this is, again, the major disadvantage of the ECB mode: Identical plaintexts are mapped to identical ciphertexts. Here is what happens in the example. For encryption, the image is broken down into small squares, which are individually encrypted. The background consists mainly of identical (white) plaintext blocks that yield a fairly uniform-looking background in the ciphertext image. On the other hand, all plain- text blocks that contain part of the (black) letters result in random-looking cipher- texts. These random-looking ciphertexts are clearly distinguishable from the uni- form background by the human eye. Fig. 5.3 Original image (top) and encrypted image using AES with 256-bit key in ECB mode (bottom) This weakness is similar to the attack on the substitution cipher that was intro- duced in Example 5.1. In both cases, statistical properties in the plaintext are pre- served in the ciphertext. Note that unlike an attack against the substitution cipher or the above banking transfer attack, an attacker does not have to do anything in the case here. The human eye automatically makes use of the statistical information. Both attacks above were examples of the weakness of a deterministic encryp- tion scheme. Thus, it is generally preferable that different ciphertexts are produced every time we encrypt the same plaintext. This behavior is called probabilistic en-
5.1 Modes of Operation for Encryption and Authentication 153 cryption. This can be achieved by introducing some form of randomness, typically in the form of an initialization vector (IV). The following modes of operation en- crypt probabilistically by means of an IV. 5.1.2 Cipher Block Chaining Mode (CBC) and Initialization Vectors There are two main ideas behind the cipher block chaining (CBC) mode. First, the encryptions of all blocks are “chained together” such that ciphertext yi depends not only on block xi but on all previous plaintext blocks as well. Second, the encryption is randomized by using an initialization vector (IV). Here are the details of the CBC mode. The ciphertext yi, which is the result of the encryption of plaintext block xi, is fed back to the cipher input and XORed with the succeeding plaintext block xi+1. This XOR sum is then encrypted, yielding the next ciphertext yi+1, which can then be used for encrypting xi+2, and so on. This process is shown on the left-hand side of Figure 5.4. For the first plaintext block x1 there is no previous ciphertext. In this case, an IV is added to the first plaintext, which also allows us to make each CBC encryption probabilistic. Note that the first ciphertext y1 depends on plaintext x1 and the IV. The second ciphertext depends on the IV, x1 and x2. The third ciphertext y3 depends on the IV and x1, x2, x3, and so on. The last ciphertext is a function of all plaintext blocks and the IV. Fig. 5.4 Encryption and decryption in CBC mode When decrypting a ciphertext block yi in CBC mode, we have to reverse the two operations we have done on the encryption side. First, we have to undo the block cipher encryption by applying the decryption function e−1. After this we have to undo the XOR operation by again XORing the correct ciphertext block. This can be expressed for general blocks yi as e−1 k (yi) = xi ⊕ yi−1. The right-hand side of Figure 5.4 shows this process. Again, if the first ciphertext block y1 is decrypted,
154 5 More About Block Ciphers the result must be XORed with the initialization vector IV to determine the plaintext block x1, i.e., x1 = IV ⊕ e−1 k (y1). The entire process of encryption and decryption can be described as follows: Definition 5.1.2 Cipher block chaining mode (CBC) Let ek(x) be a block cipher of block size b; let xi and yi be bit strings of length b; and IV be a nonce of length b. Encryption (first block): y1 = ek(x1 ⊕ IV ) Encryption (other blocks): yi = ek(xi ⊕ yi−1), i ≥ 2 Decryption (first block): x1 = e−1 k (y1) ⊕ IV Decryption (other blocks): xi = e−1 k (yi) ⊕ yi−1, i ≥ 2 We now verify the correctness of the mode, i.e., we show that the decryption actually reverses the encryption. For the decryption of the first block y1, we obtain: d(y1) = e−1 k (y1) ⊕ IV = e−1 k (ek(x1 ⊕ IV )) ⊕ IV = (x1 ⊕ IV ) ⊕ IV = x1 For the decryption of all subsequent blocks yi, i ≥ 2, we obtain: d(yi) = e−1 k (yi) ⊕ yi−1 = e−1 k (ek(xi ⊕ yi−1)) ⊕ yi−1 = (xi ⊕ yi−1) ⊕ yi−1 = xi The initialization vector IV If we choose a new IV every time we encrypt, the CBC mode becomes a probabilistic encryption scheme. If we encrypt a string of blocks x1, . . . , xt once with a first IV and a second time with a different IV, the two resulting ciphertext sequences look completely unrelated to each other to an attacker. Note that we do not have to keep the IV secret. However, in most cases, we want the IV to be a nonce, i.e., a number used only once. There are many different ways of generating and agreeing on initialization values. In the simplest case, one party choses a random number and transmits it in the clear to the other party before the actual encryption starts. Alternatively it can be a counter value that is known to Alice and Bob, and it is incremented every time a new session starts (which requires that the counter value must be stored between sessions). The IV can also be derived from values such as Alice’s and Bob’s ID number, e.g., their IP addresses, together with the current time. In order to strengthen any of these methods, we can take a value as described above, ECB-encrypt it once using the block cipher with the key known to Alice and Bob, and use the resulting ciphertext as the IV. There are some advanced attacks which also require that the IV is nonpredictable. It is instructive to discuss whether the substitution attack against the bank transfer that worked for the ECB mode is applicable to the CBC mode. If the IV is properly chosen for every wire transfer, the attack will not work at all since Oscar will not recognize any patterns in the ciphertext. For the sake of argument, let’s look at the situation if the IV is kept the same for several transfers (something that should not happen if the IV is chosen correctly). In this case, he would recognize the transfers from his account at bank A to his account at bank B. However, if he substitutes
5.1 Modes of Operation for Encryption and Authentication 155 ciphertext block 4, which is his encrypted account number, in other wire transfers going from A to B, bank B would decrypt blocks 4 and 5 to some random value. Even though money would not be redirected into Oscar’s account, it might be redi- rected to some other random account. The amount would be a random value too. This is obviously also highly undesirable for banks. This example shows that even though Oscar cannot perform specific manipulations, ciphertext alterations by him can cause random changes to the plaintext, which can have major negative conse- quences. Hence, in many, if not in most, real-world systems, encryption itself is not sufficient: We also have to protect the integrity of the message. As mentioned above, message authentication codes (MACs) and digital signatures provide mes- sage integrity. There are also modes for authenticated encryption, which provide encryption and authentication in one pass, cf. Sections 13.3.3 and 13.3.4. 5.1.3 Output Feedback Mode (OFB) In the output feedback (OFB) mode a block cipher is used to build a stream cipher encryption scheme. This scheme is shown in Figure 5.5. Note that in OFB mode the key stream is not generated bitwise but instead in a blockwise fashion. The output of the cipher gives us b key stream bits, where b is the width of the block cipher used, with which we can encrypt b plaintext bits using the XOR operation. Fig. 5.5 Encryption and decryption in OFB mode The idea behind the OFB mode is quite simple. We start by encrypting an IV with a block cipher. The cipher output gives us the first set of b key stream bits, denoted by s1. The next block of key stream bits is computed by feeding the previous cipher output back into the block cipher and encrypting it. This process is repeated as shown in Figure 5.5. The OFB mode forms a synchronous stream cipher (cf. Figure 2.3) as the key stream does not depend on the plain or ciphertext. In fact, using the OFB mode is
156 5 More About Block Ciphers quite similar to using a standard stream cipher such as Salsa20 or ChaCha. Since the OFB mode forms a stream cipher, encryption and decryption are exactly the same operation. As can be seen in the right-hand part of Figure 5.5, the receiver does not use the block cipher in decryption mode, for which we would have used the notation e−1(), to decrypt the ciphertext. This is because the actual encryption is performed by the XOR function. In order to reverse it, i.e., to decrypt the ciphertex, we simply have to perform another XOR function on the receiver side. This is in contrast to ECB and CBC mode, where the data is actually decrypted by the block cipher. Encryption and decryption using the OFB scheme can be expressed as follows. Definition 5.1.3 Output feedback mode (OFB) Let ek(x) be a block cipher of block size b; let xi, yi and si be bit strings of length b; and IV be a nonce of length b. Encryption (first block): s1 = ek(IV ) and y1 = s1 ⊕ x1 Encryption (other blocks): si = ek(si−1) and yi = si ⊕ xi, i ≥ 2 Decryption (first block): s1 = ek(IV ) and x1 = s1 ⊕ y1 Decryption (other blocks): si = ek(si−1) and xi = si ⊕ yi, i ≥ 2 As a result of the use of an IV, the OFB encryption is also nondeterministic and, thus, encrypting the same plaintext twice results in different ciphertexts. As in the case of the CBC mode, the IV should be a nonce. One advantage of the OFB mode is that the block cipher computations are independent of the plaintext. Hence, one can precompute one or several blocks si of key stream material. 5.1.4 Cipher Feedback Mode (CFB) The cipher feedback (CFB) mode also uses a block cipher as a building block for a stream cipher. It is similar to the OFB mode but instead of feeding back the key stream, the ciphertext is fed back. (A more accurate name would be “Ciphertext Feedback mode” if one follows the terminology of this book.) As in the OFB mode, the key stream is not generated bitwise but instead in a blockwise fashion. The idea behind the CFB mode is as follows: To generate the first key stream block s1, we encrypt an IV. For all subsequent key stream blocks s2, s3, . . ., we encrypt the previous ciphertext. This scheme is shown in Figure 5.6. Since the CFB mode forms a stream cipher, encryption and decryption are exactly the same operation. The CFB mode is an example of an asynchronous stream cipher (cf. Figure 2.3) since the stream cipher output is also a function of the ciphertext. The formal description of the CFB mode follows.
5.1 Modes of Operation for Encryption and Authentication 157 Fig. 5.6 Encryption and decryption in CFB mode Definition 5.1.4 Cipher feedback mode (CFB) Let ek(x) be a block cipher of block size b; let xi and yi be bit strings of length b; and IV be a nonce of length b. Encryption (first block): y1 = ek(IV ) ⊕ x1 Encryption (other blocks): yi = ek(yi−1) ⊕ xi, i ≥ 2 Decryption (first block): x1 = ek(IV ) ⊕ y1 Decryption (other blocks): xi = ek(yi−1) ⊕ yi, i ≥ 2 As a result of the use of an IV, the CFB encryption is also nondeterministic; hence, encrypting the same plaintext twice results in different ciphertexts. As in the case of the CBC and OFB modes, the IV should be a nonce. A variant of the CFB mode can be used in situations where short plaintext blocks are to be encrypted. Let’s use the encryption of the link between a (remote) key- board and a computer as an example. The plaintexts generated by the keyboard are typically only 1 byte long, e.g., an ASCII character. In this case, only 8 bits of the key stream are used for encryption (it does not matter which ones we choose as they are all secure), and the ciphertext also only consists of 1 byte. The feedback of the ciphertext as input to the block cipher is a bit tricky and works as follows. The previous block cipher input is shifted by 8 bit positions to the left, and the 8 least significant positions of the input register are filled with the ciphertext byte. This process repeats. Of course, this approach works not only for plaintext blocks of length 8 but for any lengths shorter than the cipher output. 5.1.5 Counter Mode (CTR) Another mode that uses a block cipher as a stream cipher is the counter (CTR) mode. As in the OFB and CFB modes, the key stream is computed in a blockwise fashion. The input to the block cipher is a counter, which assumes a different value
158 5 More About Block Ciphers every time the block cipher computes a new key stream block. Figure 5.7 shows the principle. Fig. 5.7 Encryption and decryption in counter mode We have to be careful how we initialize the input to the block cipher. We must prevent use of the same input value twice. Otherwise, if an attacker knows one of the two plaintexts that were encrypted with the same input, he can compute the key stream block and thus immediately decrypt the other ciphertext. In order to achieve this uniqueness,the following approach is often taken in practice. Let’s assume a block cipher with an input width of 128 bits such as AES. First we choose an IV that is a nonce with a length smaller than the block length, e.g., 96 bits. The remaining 32 bits are then used by a counter with the value CTR, which is initialized to zero. For every block that is encrypted during the session, the counter is incremented but the IV stays the same. In this example, the number of blocks we can encrypt without choosing a new IV is 232. Since every block consists of 16 bytes, a maximum of 16 × 232 = 236 bytes, or about 64 Gigabytes, can be encrypted before a new IV must be generated. Here is a formal description of the counter mode with a cipher input construction as just introduced. Definition 5.1.5 Counter mode (CTR) Let ek(x) be a block cipher of block size b, and let xi and yi be bit strings of length b. The concatenation of the initialization value IV and the counter CTRi is denoted by (IV ||CTRi) and is a bit string of length b. Encryption: yi = ek(IV ||CTRi) ⊕ xi, i ≥ 1 Decryption: xi = ek(IV ||CTRi) ⊕ yi, i ≥ 1 Please note that the string (IV ||CTR1) does not have to be kept secret. It can, for instance, be generated by Alice and sent to Bob together with the first ciphertext
5.1 Modes of Operation for Encryption and Authentication 159 block. The counter CTR can either be a regular integer counter or a slightly more complex function such as a maximum-length LFSR. One attractive feature of the counter mode is that it can be parallelized because, unlike the OFB or CFB mode, it does not require any feedback. For instance, we can have two block cipher engines running in parallel, where the first block cipher encrypts the counter value CTR1 and the other CTR2 at the same time. When the two block cipher engines are finished, the first engine encrypts the value CTR3 and the other one CTR4, and so on. This scheme would allow us to encrypt at twice the data rate of a single implementation. Of course, we can have more than two block ciphers running in parallel, increasing the speed-up accordingly. For applications with high throughput demands, e.g., in networks with data rates in the range of Gigabits per second, encryption modes that can be parallelized are often desirable. 5.1.6 XTS-AES Unlike the modes of operation introduced so far, the XTS-AES mode was designed for one specific application scenario: XTS-AES is optimized for the encryption of data on storage devices such as hard disks and uses the AES block cipher as a building block. The acronym XTS stands for the XEX Tweakable Block Cipher with Ciphertext Stealing where XEX is the short term for XOR-Encrypt-XOR. This mode for storage encryption is specifically designed to enable random and independent access to encrypted data blocks on the storage device. With any of the chaining- based modes of operation that we have seen earlier in this chapter, it is clearly not possible to just encrypt or decrypt individual blocks. The mode encrypts a data stream divided into consecutive equally sized data units, which are then stored on the storage device. The length of the data unit is typically based on the block or sector size of the storage device. The data unit should have a minimum length of 128 bits. The core idea behind this mode is that it upgrades the AES block cipher into a so-called tweakable block cipher. The tweak can be considered as another input to the block cipher in addition to the plaintext data and the secret key. That way the block cipher can accept another 128-bit input that incorporates the logical position of the data unit (i.e., the physical address of the data block) on the storage device in the encryption process. With this approach, even two identical plaintexts that are stored at different positions on the storage device result in two (completely) different ciphertexts. This prevents an adversary from gaining any information from the ciphertext. Since the encryption process depends only on the tweak and thus the physical address of the encrypted data, it allows for parallelization and pipelining in implementations. The encryption procedure of one data unit is shown in Figure 5.8 and is defined as follows.
160 5 More About Block Ciphers Definition 5.1.6 XTS-AES Encryption Given is the AES block cipher with a key length of 128 or 256 bits. The XTS key k is a concatenation of two equally sized AES subkeys k1 and k2, i.e., k = (k1||k2), where k has 256 or 512 bits. Let x be a data unit consisting of plaintext blocks x1, . . . , xn, and corresponding ciphertexts y1, . . . , yn. Each 128-bit plaintext block is assigned a counter value j = 1, . . . , n per data unit. Finally we denote with i the 128-bit tweak value that determines the location (i.e., physical address) of the data unit on the storage device. Encryption 1. T = AESk2 (i) ⊗ α j 2. y j = AESk1 (x j ⊕ T ) ⊕ Tjk2 k1 j j AES AES Fig. 5.8 XTS-AES encryption of one data unit The mode uses a multiplication in the Galois Field GF(2128) to compute a mask T . The mask is XORed to the plaintext and the ciphertext and depends both on the tweak i and on the counter j. For computing T , the 16-byte output of the en- crypted tweak AESk2 (i) is represented as an element of GF(2128) and multiplied by α j modulo the irreducible polynomial x128 + x7 + x2 + x + 1, where α = x is a prim- itive element of the Galois field. We note that for a given data unit, the tweak i is only encrypted once but a different masking value T is computed for every 128-bit plaintext block x j through the varying value of j. Decryption in the XTS-AES mode is very similar to encryption. If the data unit is not a multiple of 128 bits, so-called ciphertext stealing is used to pad the last plaintext block [145].
5.2 Exhaustive Key Search Revisited 161 5.2 Exhaustive Key Search Revisited In Section 3.5.1 we saw that given a plaintext–ciphertext pair (x1, y1) a DES key can be found by exhaustive search using the simple algorithm: DESki (x1) ? = y1, i = 0, 1, . . . , 256 − 1 (5.1) In practice, however, a key search is often more complicated. Somewhat surpris- ingly, a brute-force attack can produce false positive results, i.e., keys ki are found that are not the one used for the encryption by Alice, yet they perform a correct en- cryption in Equation (5.1). The likelihood of this occurring is related to the relative size of the key space and the plaintext space. In order to find the correct key several pairs of plaintext–ciphertext are needed. The length of the respective plaintext required to break the cipher with a brute-force attack is referred to as unicity distance. Let us first look why one pair (x1, y1) might not be sufficient to identify the cor- rect key. For illustration purposes we assume a cipher with a block width of 64 bits and a key size of 80 bits (PRESENT is a cipher with such parameters). If we en- crypt x1 under all possible 280 keys, we obtain 280 ciphertexts. However, there exist only 264 different ones, and thus some keys must map x1 to identical ciphertexts. If we run through all keys for a given plaintext-ciphertext pair, we find on average 280/264 = 216 keys that perform the mapping ek(x1) = y1. This estimation is valid since the encryption of a plaintext for a given key can be viewed as a random se- lection of a 64-bit ciphertext string. The phenomenon of multiple “paths” between a given plaintext and ciphertext is depicted in Figure 5.9, in which k(i) denote the keys that map x1 to y1. These keys can be considered key candidates. Fig. 5.9 Multiple keys map between one plaintext and one ciphertext Among the approximately 216 key candidates k(i) is the correct one that was used by Alice to perform the encryption. Let’s call this one the target key. In order to
162 5 More About Block Ciphers identify the target key we need a second plaintext–ciphertext pair (x2, y2). Again, there are about 216 key candidates that map x2 to y2. One of them is the target key. The other keys can be viewed as randomly drawn from the 280 possible ones. It is crucial to note that the target key must be present in both sets of key candidates. To determine the effectiveness of a brute-force attack, the crucial question is now: What is the likelihood that another (false!) key is contained in both sets? The answer is given by the following theorem. Theorem 5.2.1 Given a block cipher with a key length of κ bits and block size of n bits, as well as t plaintext–ciphertext pairs (x1, y1), . . . , (xt , yt ), the expected number of false keys which en- crypt all plaintexts to the corresponding ciphertexts is: 2κ−tn Returning to our example and assuming two plaintext–ciphertext pairs, the likeli- hood of a false key k f that performs both encryptions ek f (x1) = y1 and ek f (x2) = y2 is: 280−2·64 = 2−48 This value is so small that for almost all practical purposes it is sufficient to test two plaintext–ciphertext pairs. If the attacker chooses to test three pairs, the likelihood of a false key decreases to 280−3·64 = 2−112. As we see from this example, the like- lihood of a false alarm decreases rapidly with the number t of plaintext–ciphertext pairs. In practice, typically we only need a few pairs. The theorem above is not only important if we consider an individual block ci- pher but also if we perform multiple encryptions with a cipher. This issue is ad- dressed in the following section. 5.3 Increasing the Security of Block Ciphers In some situations we wish to increase the security of block ciphers, e.g., if a cipher such as DES is available in hardware or software for legacy reasons in a given application. We discuss two general approaches to strengthen a cipher: multiple encryption and key whitening. Multiple encryption, i.e., encrypting a plaintext more than once, is already a fundamental design principle of block ciphers, since the round function is applied many times to the cipher. Our intuition tells us that the security of a block cipher against both brute-force and analytical attacks increases by performing multiple encryptions in a row. Even though this is true in principle, there are a few surprising facts. For instance, doing double encryption does very little to increase the brute-force resistance over a single encryption if the attacker has a lot of storage space available. We study this counterintuitive fact in the next section.
5.3 Increasing the Security of Block Ciphers 163 Another very simple yet effective approach to increase the brute-force resistance of block ciphers is called key whitening; it is also discussed below. We note that when using AES, we already have three different security levels given by the key lengths of 128, 192 and 256 bits. Since there are no realistic at- tacks known against AES with any of those key lengths, there appears no reason to perform multiple encryption with AES for practical systems. However, for some selected older ciphers, especially for DES, multiple encryption can be a useful tool. 5.3.1 Double Encryption and Meet-in-the-Middle Attack Let us assume a block cipher with a key length of κ bits. For double encryption, a plaintext x is first encrypted with a key kL, and the resulting ciphertext is encrypted again using a second key kR. This scheme is shown in Figure 5.10. Fig. 5.10 Double encryption and meet-in-the-middle attack A na¨ıve brute-force attack would require searching through all possible combi- nations of both keys, i.e., the effective key length would be 2κ and an exhaustive key search would require 2κ ·2κ = 22κ encryptions (or decryptions). However, using the meet-in-the-middle attack, the key space is drastically reduced. This is a divide- and-conquer attack in which Oscar first brute-force attacks the encryption on the left-hand side, which requires 2κ cipher operations, and then the right encryption, which again requires 2κ operations. If he succeeds with this attack, the total com- plexity is 2κ + 2κ = 2 · 2κ = 2κ+1. This is only twice as costly as a key search of a single encryption and of course dramatically less complex than performing 22κ search operations. The attack has two phases. In the first one, the left encryption is brute-forced and a lookup table is computed. In the second phase the attacker tries to find a match in
164 5 More About Block Ciphers the table, which reveals both encryption keys. Here are the details of the meet-in- the-middle attack. Phase I: Table Computation For a given plaintext x1, compute a lookup table for all pairs (kL,i, zL,i), where ekL,i (x1) = zL,i and i = 1, 2, . . . , 2κ . These computations are symbolized by the left arrow in the figure. The zL,i are the intermediate values that occur in between the two encryptions. This list should be ordered by the values of the zL,i. The number of entries in the table is 2κ , with each entry being n + κ bits wide. Note that one of the keys we used during the table construction must be the correct target key, but we still do not know which one it is. Phase II: Key Matching In order to find the target key, we now decrypt y1, i.e., we perform the computations symbolized by the right arrow in the figure. We select the first possible key kR,1, e.g., the all-zero key, and compute: e−1 kR,1 (y1) = zR,1 We now check whether zR,1 is equal to any of the zL,i values in the table which we computed in the first phase. If it is not in the table, we increment the key to kR,2, decrypt y1 again, and check whether this value is in the table. We continue until we have a match. Such a match is also called a collision of two values, i.e., zL,i = zR, j. This gives us two keys: The value zL,i is associated with the key kL,i from the left encryption, and kR, j is the key we just tested in the decryption coming from the right side. This means there exists a key pair (kL,i, kR, j) which performs the double encryption: ekR, j (ekL,i (x1)) = y1 (5.2) As discussed in Section 5.2, there is a chance that this is not the target key pair we are looking for if there are several possible key pairs that perform the mapping x1 → y1. Hence, we have to verify additional key candidates by encrypting several plaintext–ciphertext pairs according to Equation (5.2). If the verification fails for any of the pairs (x1, y1), (x2, y2), . . ., we go back to beginning of Phase II and increment the key kR again and continue with the search. Let us briefly discuss how many plaintext–ciphertext pairs we will need to rule out faulty keys with a high likelihood. With respect to multiple mappings between a plaintext and a ciphertext as depicted in Figure 5.9, double encryption can be modeled as a cipher with 2κ key bits and n block bits. In practice, one often has 2κ > n, in which case we need several plaintext–ciphertext pairs. Theorem 5.2.1 can easily be adopted to the case of multiple encryption, which gives us a useful guideline about how many (x, y) pairs should be available.
5.3 Increasing the Security of Block Ciphers 165 Theorem 5.3.1 Given are l subsequent encryptions with a block cipher with a key length of κ bits and block size of n bits, as well as t plaintext–ciphertext pairs (x1, y1), . . . , (xt , yt ). The expected num- ber of false keys which encrypt all plaintexts to the corresponding ciphertexts is given by: 2lκ−tn Let us look at an example. Example 5.3. As an example, if we double-encrypt with DES and choose to test three plaintext–ciphertext pairs, the likelihood of a faulty key pair surviving all three key tests is: 22·56−3·64 = 2−80 Let us examine the computational complexity of the meet-in-the-middle attack. In the first phase of the attack, corresponding to the left arrow in Figure 5.10, we perform 2κ encryptions and store them in 2κ memory locations. In the second stage, corresponding to the right arrow in the figure, we perform a maximum of 2κ decryp- tions and table look-ups. We ignore multiple-key tests at this stage. The total cost for the meet-in-the-middle attack is: number of encryptions and decryptions = 2κ + 2κ = 2κ+1 number of storage locations = 2κ This compares to 2κ encryptions or decryptions and essentially no storage cost in the case of a brute-force attack against a single encryption. Even though the storage requirements go up quite a bit, the costs in computation and memory are still only proportional to 2κ . Thus, it is widely believed that double encryption is not worth the effort. Instead, triple encryption should be used; this method is described in the following section. Note that for a more exact complexity analysis of the meet-in-the-middle attack, we would also need take the cost of sorting the table entries in Phase I into account as well as the table look-ups in Phase II. For our purposes, however, we can ignore these additional costs. 5.3.2 Triple Encryption Compared to double encryption, a much more secure approach is the encryption of a block of data three times in a row: y = ek3 (ek2 (ek1 (x)))
166 5 More About Block Ciphers In practice, often the following version of triple encryption is used: y = ek1 (e−1 k2 (ek3 (x))) This type of triple encryption is sometimes referred to as encryption–decryption– encryption (EDE). The reason for this has nothing to do with security. If k1 = k2, the operation effectively performed is y = ek3 (x) which is single encryption. Since it is sometimes desirable that one implementation can perform both triple encryption and single encryption, e.g., in order to interop- erate with legacy systems, EDE is a popular choice for triple encryption. Triple encryption is in practice especially relevant in the case of DES, referred to as 3DES or triple DES and described in Section 3.7.2. Of course, we can still perform a meet-in-the-middle attack as shown in Fig- ure 5.11. Again, we assume κ bits per key. The problem for an attacker is that she Fig. 5.11 Triple encryption and sketch of a meet-in-the-middle attack has to compute a lookup table either after the first or after the second encryption. In both cases, the attacker has to compute two encryptions or decryptions in a row in order to reach the lookup table. Here lies the cryptographic strength of triple encryp- tion: There are 22k possibilities to run through all possible keys of two encryptions or decryptions. In the case of 3DES, this forces an attacker to perform 2112 key tests, which is out of reach with current technology. In summary, the meet-in-the-middle attack reduces the effective key length of triple encryption from 3 κ to 2 κ. Because of this, it is often said that the effective key length of 3DES is 112 bits as opposed to the 3 · 56 = 168 bits that are actually used as key input to the cipher.
5.3 Increasing the Security of Block Ciphers 167 5.3.3 Key Whitening Using an extremely simple technique called key whitening, it is possible to make block ciphers such as DES much more resistant against brute-force attacks. The basic scheme is shown in Figure 5.12. Fig. 5.12 Key whitening of a block cipher In addition to the regular cipher key k, two whitening keys k1 and k2 are used to XOR-mask the plaintext and ciphertext. This process can be expressed as follows. Definition 5.3.1 Key whitening for block ciphers Encryption: y = ek,k1,k2 (x) = ek(x ⊕ k1) ⊕ k2 Decryption: x = e−1 k,k1,k2 (y) = e−1 k (y ⊕ k2) ⊕ k1 It is important to stress that key whitening does not strengthen block ciphers against most analytical attacks such as linear and differential cryptanalysis. This is in contrast to multiple encryption, which often also increases the resistance to analytical attacks. Hence, key whitening is not a “cure” for inherently weak ciphers. Its main application is to ciphers that are strong against analytical attacks but possess a key space that is too short. The prime example of such a cipher is DES. A variant of DES which uses key whitening is DESX. In the case of DESX, the key k2 is derived from k and k1. Please note that most modern block ciphers such as AES already apply key whitening internally by adding a subkey prior to the first round and after the last round. Let’s now discuss the security of key whitening. A na¨ıve brute-force attack against the scheme requires 2κ+2n search steps, where κ is the bit length of the key and n the block size. Using the meet-in-the-middle attack introduced in Sec- tion 5.3.1, the computational load can be reduced to approximately 2κ+n steps, plus storage of 2n data sets. However, if the adversary Oscar can collect 2m plaintext– ciphertext pairs, a more advanced attack exists with a computational complexity of 2κ+n−m
168 5 More About Block Ciphers cipher operations. Even though we do not introduce the attack here, we’ll briefly discuss its consequences if we apply key whitening to DES. We assume that the attacker knows 2m plaintext–ciphertext pairs. Note that the designer of a security system can often control how many plaintext–ciphertext pairs are generated before a new key is established. Thus, the parameter m often cannot be arbitrarily increased by the attacker. Also, since the number of known plaintexts grows exponentially with m, values beyond, say, m = 40, seem quite unrealistic. As a practical example, let’s assume key whitening of DES, and that Oscar can collect a maximum of 232 plaintexts (that is about 34 GB of data). He now has to perform 256+64−32 = 288 DES computations. It can be speculated that 288 encryptions is at the edge of what large government agencies can do. Thus, even though key-whitening is useful for such a surprisingly simple technique, it does not provide long-term security if used together with DES. 5.4 Discussion and Further Reading NIST’s Modes of Operation After the AES selection, the U.S. National Institute of Standards and Technology (NIST) supported the process of evaluating new modes of operation in a series of special publications and workshops [194]. Table 5.1 pro- vides an overview of the most relevant modes of operation standardized by NIST. The NIST Special Publications 800-38 A–F describe five modes for confidentiality Table 5.1 NIST modes of operations in Special Publications 800-38 SP 800-38 A B C D E F [101] [105] [103] [102] [104] [106] modes ECB, CBC, CFB, OFB, CTR CMAC CCM GCM, GMAC XTS-AES KW, KWP, TKW confidentiality X X (X) X authentication X X X key wrap X of data (ECB, CBC, CFB, OFB, CTR), one mode for storage encryption (XTS- AES), one for authentication (CMAC), two combined modes for confidentiality and authentication (CCM, GCM), and three modes for key wrapping (KW, KWP, TKW). Modes CMAC, CCM and GCM are described in Chapter 13. The NIST modes are widely used in practice and are part of many industry standards, e.g., for computer networks or banking. We note that the NIST XTS-AES standard refers to the IEEE
5.4 Discussion and Further Reading 169 Standard 1619-2018 [145] and is based on the XEX (XOR Encrypt XOR) tweakable block cipher, proposed by Phillip Rogaway [220]. The original XEX proposal can only encrypt messages that are exact multiples of 128-bit blocks, whereas NIST’s XTS-AES mode allows input data with arbitrary length. Key wrapping modes (KW, KWP, TKW) Key wrapping describe methods to pro- tect the confidentiality as well as the authenticity and integrity of cryptographic keys. The AES Key Wrap (KW) mode is a deterministic authenticated encryption mode of operation. Its variant with an internal padding scheme is called AES Key Wrap With Padding (KWP). The main component of the key wrap modes is a block cipher. The key for the underlying block cipher is called the key encryption key (KEK), denoted K. For the KW and KWP modes, the underlying block cipher shall have a key size of 128 bits or more, which is of course the case for AES. For TKW, the underlying block cipher is specified to be 3DES, and the block size is therefore 64 bits. The KEK has to be kept secret and its key length directly affects the security of the algorithms against brute-force attack. For a full description of the AES Key Wrap (KW) and its variants KWP and TKW we refer to the NIST document [106]. Authenticated Encryption and Other Applications for Block Ciphers The most important application of block ciphers in practice, in addition to data encryption, is Message Authentication Codes (MACs), which are discussed in Chapter 13. In addi- tion to the CMAC message authentication mode, which is listed in the table above and described in Section 13.3.2, there are several other MACs that are constructed with a block cipher, including OMAC and PMAC. Authenticated Encryption (AE) uses block ciphers to encrypt and generate a MAC at the same time in order to provide both confidentiality and authentication. The table above contains two such modes, CCM and GCM, which are described in Sections 13.3.3 and 13.3.4. Other authenticated encryption modes include the EAX mode, OCB mode and GC mode. Another application for block ciphers is the Cryptographically Secure Pseudo Random Number Generator (CSPRNG). In fact, the stream cipher modes introduced in this chapter, OFB, CFB and CTR mode, form CSPRNGs. There are also standards such as [13, Appendix A.2.4] which explicitly specify random number generators from block ciphers. Block ciphers can also be used to build cryptographic hash functions, as dis- cussed in Chapter 11. Brute-Force and Quantum Computer Attacks Even though there are no algo- rithmic shortcuts to brute-force attacks, there are methods that are efficient if sev- eral exhaustive key searches have to be performed. Those methods are called time– memory tradeoff attacks (TMTO). The general idea is to encrypt a fixed plaintext under a large number of keys and to store certain intermediate results. This is the precomputation phase, which is typically at least as complex as a single brute-force attack and which results in large lookup tables. In the online phase, a search through the tables takes place that is considerably faster than a brute-force attack. Thus, after
170 5 More About Block Ciphers the precomputation phase, individual keys can be found much more quickly. TMTO attacks were originally proposed by Martin Hellman [141] and were improved with the introduction of distinguished points by Ronald Rivest [218]. Later, rainbow ta- bles were proposed by Philippe Oechslin to further improve TMTO attacks [206]. A limiting factor of TMTO attacks in practice is that for each individual attack it is required that the same piece of known plaintext was encrypted, e.g., a file header. Attacking block ciphers (or stream ciphers) with quantum computers which might become available in the future is discussed in Section 12.1. The main obser- vation is that a potential quantum computer would require only 2n/2 steps in order to perform a key search on a cipher with an n-bit key, using Grover’s algorithm [128]. 5.5 Lessons Learned There are many different ways to encrypt with a block cipher. The modes of operation have their specific advantages and disadvantages. Some modes of operation turn a block cipher into a stream cipher. The straightforward ECB mode has security weaknesses, independent of the un- derlying block cipher. The counter mode allows parallelization of encryption and is thus suited for high- speed implementations. Double encryption with a given block cipher only marginally improves the resis- tance against brute-force attacks. Triple encryption with a given block cipher roughly doubles the effective key length. For example, triple DES (3DES) has an effective key length of 112 bits.
5.5 Problems 171 Problems 5.1. Assume a toy block cipher e() for encryption of 5-bit blocks. The encryption function is a bit permutation, which depends on the key. We assume that for a given key the encryption (permutation) is as follows: e(b1b2b3b4b5) = (b2b5b4b1b3) Encrypt the message x = 01101 11011 11010 00110 with the five different modes of operation ECB, CBC, CFB, OFB and CTR, and provide the corresponding cipher- text y. Use IV = 11001 as initialization vector. 5.2. We consider exhaustive key-search attacks on block ciphers where the key is k bits long. The block length is n bits, with n being much larger than k. 1. Exhaustive key searches typically require known plaintexts. How many plaintext– ciphertext pairs are needed to successfully break the block cipher running in ECB mode? How many steps are needed in the worst case? 2. Assume that the initialization vector IV for running the block cipher in CBC mode is known (which is in practice the case as the IV is transmitted unen- crypted). How many plaintext–ciphertext pairs are now needed to break the ci- pher by performing an exhaustive key search? How many steps are needed max- imally? Briefly describe the attack. 3. How many plaintext–ciphertext pairs are necessary if you do not know the IV? 4. Is breaking a block cipher in CBC mode by means of an exhaustive key search considerably more difficult than breaking an ECB-mode block cipher? 5.3. In a company, all files that are sent on the internal network are automatically encrypted by using AES-128 in CBC mode. A fixed key is used, and the IV is changed once per day. The network encryption is file-based, so that the IV is used at the beginning of every file. Through hacking into the system you manage to find the fixed AES-128 key but you do not know the current IV. Today, you were able to eavesdrop and obtain two different files, one with unknown content and one which is known to be an automatically generated temporary file, which only contains the value 0xFF. Briefly describe how it is possible to obtain the unknown initialization vector and how you are able to decrypt the unknown file. 5.4. Keeping the IV secret in OFB mode does not make an exhaustive key search more complex. Describe how we can perform a brute-force attack with unknown IV. What are the requirements regarding plaintext and ciphertext? 5.5. Describe how the OFB mode can be attacked if the IV is not different for each execution of the encryption operation. 5.6. Propose a simple change to the OFB mode that encrypts one byte of plaintext at a time, e.g., for encrypting key strokes from a remote keyboard. The block cipher used is AES. Perform one block cipher operation for every new plaintext byte. Draw a block diagram of your scheme and pay particular attention to the bit lengths used in your diagram.
172 5 More About Block Ciphers 5.7. As is so often true in cryptography, it is easy to weaken a seemingly strong scheme by small modifications. Assume a variant of the OFB mode by which we only feed back the 8 most significant bits of the cipher output. We use AES and fill the remaining 120 input bits of the cipher with 0s. 1. Why is this scheme weak if we encrypt moderately large blocks of plaintext, say 100 kBytes? What is the maximum number of known plaintexts an attacker needs to completely break the scheme? 2. Let the feedback byte be denoted by FB. Does the scheme become cryptographi- cally stronger if we feed back the 128-bit value FB, FB, . . . , FB to the input (i.e., we copy the feedback byte 16 times and use it as AES input)? 5.8. In the text, a variant of the CFB mode is proposed that encrypts individual bytes. Draw a block diagram for this mode when using AES as block cipher. Indicate the width (in bits) of each line in your diagram. 5.9. We are using AES in counter mode to encrypt a hard disk with 1 TB of capacity. What is the maximum length of the IV? 5.10. Sometimes error propagation is an issue when choosing a mode of operation in practice. In order to analyze the propagation of errors, let us assume a bit error (i.e., a substitution of a “0” bit by a “1” bit or vice versa) in a ciphertext block yi. Alice is sending messages to Bob. 1. Assume an error occurs during the transmission in one block of ciphertext, let’s say yi. Which plaintext blocks are affected on Bob’s side when using the ECB mode? 2. Again, assume block yi contains an error introduced during transmission. Which plaintext blocks are affected on Bob’s side when using the CBC mode? 3. Suppose there is an error in the plaintext xi on Alice’s side. Which plaintext blocks are affected on Bob’s side when using the CBC mode? 4. Assume a single-bit error occurs in the transmission of a ciphertext character in 8-bit CFB mode. How far does the error propagate? Describe exactly how each block is affected. 5. Give an overview of the effect of bit errors in a ciphertext block for the modes ECB, CBC, CFB, OFB and CTR. Differentiate between random bit errors and specific bit errors when decrypting yi. Specific bit errors means errors at the same position(s) as the original bit error(s). 5.11. Besides simple bit errors, the deletion or insertion of a bit during transmission can yield even more severe effects for many modes of operation since the synchro- nization of blocks is disrupted. In most cases, the decryption of subsequent blocks will be incorrect. A special case is the CFB mode with a feedback width of 1 bit. Show that the synchronization is automatically restored after κ + 1 steps, where κ is the block size of the block cipher. 5.12. We now analyze the security of DES with double encryption (2DES) by doing a cost estimate. The encryption is described by the following expression:
5.5 Problems 173 2DES(x) = DESK2 (DESK1 (x)) 1. First, let us assume a key search without building lookup tables. For this purpose, the whole key space spanned by K1 and K2 has to be searched. How much does a key-search machine for breaking 2DES (worst case) in 1 week cost? We assume we have ASICs that can test 107 keys per second at a cost of $5 per IC. Furthermore, assume an overhead of 50% for building the key-search machine. 2. Let us now consider the meet-in-the-middle (or time-memory tradeoff) attack that was introduced in this chapter, in which we can use lookup tables. Answer the following questions: How many entries have to be stored? How many bytes (not bits!) have to be stored for each entry? How costly is a key search in one week? Please note that the key space has to be searched before filling up the memory completely. Then we can begin to search the key space of the second key. Assume the same hardware for both key spaces. For a rough cost estimate, assume the following costs for hard disk space: $5/1 TByte, where 1 TByte = 1012 Bytes. 3. Assuming that both processing costs and the price for storage decrease according to Moore’s law, i.e., they decrease by 50% every 18 months, when do the total costs move below $1 million? 5.13. Let e be a block cipher with a block size of n = 64 bits and a key length of κ = 80 bits. A practical example of such a cipher is PRESENT. We assume the algorithm does not have any mathematical weaknesses that can be exploited. Let us now look at the following figure showing a triple encryption scheme: 1. Provide an expression for the encryption and decryption. 2. Why does this scheme use the encryption-decryption-encryption (EDE) mode (as opposed to encrypting three times)? Does this choice increase the security compared to triple encryption?
174 5 More About Block Ciphers 3. How complex is a simple brute-force attack on this cipher in EDE mode? Com- pare the computational complexity and the storage requirements with the com- plexity of a meet-in-the-middle (MITM) attack. What is the effective key length of this triple encryption? Do you think that either of these attacks is feasible with today’s computers? 4. How many blocks of plaintext are required for the MITM attack in order find the correct key with high probability? 5.14. Imagine that aliens — rather than abducting earthlings and performing strange experiments on them — drop a computer on planet Earth that is particularly suited for AES key searches. In fact, it is so powerful that we can search through 128, 192 and 256 key bits in a matter of days. Provide guidelines for the number of plaintext– ciphertext pairs the aliens need so that they can rule out false keys with a reasonable likelihood. (Remark: Since the existence of both aliens and human-built computers for such key lengths seem extremely unlikely at the time of writing, this problem is pure science fiction.) 5.15. In this problem, you are going to attack a symmetric multiple encryption scheme with given pairs of plain- and ciphertexts. 1. You want to break an encryption scheme using triple AES-192 encryption with a block length of n = 128 bits and a key length of k = 192 bits. How many plaintext/ciphertext pairs are required to reduce the probability of obtaining a wrong key candidate K′ in a brute-force attack to at most Pr(K′ 6 = K) = 2−20? 2. Assume a block cipher with a block length of n = 80 and a double encryption scheme (l = 2). What is the maximum key length of the block cipher such that you can attack the cipher with an effective error rate of Pr(K′ 6 = K) = 2−10 ≈ 1/1024? 3. Estimate the probability of success in case of a double encryption with AES-256 (i.e., a block length of n = 128 bits and a key length of k = 256 bits) with four given ciphertext pairs at hand? 5.16. 3DES with three different keys can be broken with about 22k encryptions and 2k memory cells, where k = 56. Design the corresponding attack. How many pairs (x, y) should be available so that the probability of determining an incorrect key triple (k1, k2, k3) is sufficiently low? 5.17. This is your chance to break a cryptosystem. As we know by now, cryptogra- phy is a tricky business. The following problem illustrates how easy it is to turn a strong scheme into a weak one with minor modifications. We saw in this chapter that key whitening is a good technique for strengthening block ciphers against brute-force attacks. We now look at the following variant of key whitening against DES, which we’ll call DESA: DESAk,k1 (x) = DESk(x) ⊕ k1 Even though the method looks similar to key whitening, it hardly adds to the se- curity. Your task is to show that breaking the scheme is roughly as difficult as a
5.5 Problems 175 brute-force attack against single DES. Assume you have a few plaintext-?ciphertext pairs. 5.18. Let us now consider a brute-force attack on a block cipher with key length k. The block cipher is used in OFB mode. The initialization vector is not known. Describe how many (i) plaintexts and (ii) ciphertexts are required to break the cipher with a brute-force attack. In the worst case, how many steps are necessary? 5.19. Draw a diagram of the decryption process of the XTR-AES mode. It is suffi- cient to show the encryption of one data unit, in analogy to Figure 5.8.
Chapter 6 Introduction to Public-Key Cryptography As we start to learn about public-key cryptography, let’s recall that the term public- key cryptography is used interchangeably with asymmetric cryptography; they both denote exactly the same thing and are used synonymously. As stated in Chapter 1, symmetric cryptography has been used for at least 4000 years. Public-key cryptography, on the other hand, is quite new. It was publicly introduced by Whitfield Diffie, Martin Hellman and Ralph Merkle in 1976. Two decades later, in 1997, British documents that were declassified revealed that the re- searchers James Ellis, Clifford Cocks and Graham Williamson from UK’s Govern- ment Communications Headquarters (GCHQ) discovered and realized the principle of public-key cryptography a few years earlier, in 1972. However, it is still being de- bated whether GCHQ fully recognized the far-reaching consequences of public-key cryptography for commercial security applications. In this chapter you will learn: A brief history of public-key cryptography The pros and cons of public-key cryptography Some number-theoretical topics that are needed for understanding public-key algorithms, most importantly the extended Euclidean algorithm 177 C. Paar et al., Understanding Cryptography, https://doi.org/10.1007/978-3-662-69007-9_6 © The Editor(s) (if applicable) and The Author(s), under exclusive license to Springer-Verlag GmbH, DE, part of Springer Nature 2024
178 6 Introduction to Public-Key Cryptography 6.1 Symmetric vs. Asymmetric Cryptography In this chapter we will see that asymmetric, i.e., public-key, algorithms are very dif- ferent from symmetric algorithms such as AES or DES. Most public-key algorithms are based on number-theoretic functions. This is quite different from symmetric ci- phers, where the goal is usually not to have a compact mathematical description between input and output. Even though mathematical structures are often used for small blocks within symmetric ciphers, for instance, in the AES S-box, this does not mean that the entire cipher has a compact mathematical description. Symmetric Cryptography Revisited In order to understand the principle of asymmetric cryptography, let us first recall the basic symmetric encryption scheme in Figure 6.1.BobAlice Fig. 6.1 Principle of symmetric-key encryption Such a system is symmetric with respect to two properties: 1. The same secret key is used for encryption and decryption. 2. The encryption and decryption functions are very similar (in the case of DES they are essentially identical). There is a simple analogy for symmetric cryptography, as shown in Figure 6.2. Assume there is a safe with a strong lock. Only Alice and Bob have a copy of the key for the lock. The action of encrypting a message can be viewed as putting the message in the safe. In order to read, i.e., decrypt, the message, Bob uses his key and opens the safe. Modern symmetric algorithms such as AES or 3DES are very secure, fast and are in widespread use. However, there are several shortcomings associated with symmetric-key schemes, as discussed below. Key Distribution Problem The key must be established between Alice and Bob using a secure channel. Remember that the communication link for the message is
6.1 Symmetric vs. Asymmetric Cryptography 179BobAlice Fig. 6.2 Analogy for symmetric encryption: a safe for which both Alice and Bob have a key not secure, so sending the key over the channel directly — which would be the most convenient way of transporting it — can not be done. Number of Keys Even if we solve the key distribution problem, we must poten- tially deal with a very large number of keys. If each pair of users needs a separate key for secure communication in a network with n users, there are n · (n − 1) 2 keys, and every user has to store n − 1 keys securely. Even for mid-size networks, say, a corporation with 2000 people, this requires approx. 2 million keys that must be generated and transported via secure channels. More about this problem is found in Section 14.1. (There are smarter ways of dealing with keys in symmetric cryptog- raphy networks, as detailed in Section 14.3; however, those approaches have other problems such as a single point of failure.) No Protection Against Cheating by Alice or Bob Alice and Bob have the same capabilities, since they possess the same key. As a consequence, symmetric cryp- tography cannot be used in situations where we would like to prevent cheating by either Alice or Bob as opposed to cheating by an outsider like Oscar. For instance, in e-commerce applications it is often important to prove that Alice actually sent a certain message, say, an online order for a flat screen TV. If we only use symmetric cryptography and Alice changes her mind later, she can always claim that Bob, the vendor, has falsely generated the electronic purchase order. Preventing this is called non-repudiation and can be achieved with asymmetric cryptography, as discussed in Section 10.1.1, or with digital signatures, which we introduce in Chapter 10.
180 6 Introduction to Public-Key CryptographyBobAlice public key private key unlockdeposit Fig. 6.3 Analogy for public-key encryption: a safe with a public lock for depositing a message and a secret lock for retrieving a message Principles of Asymmetric Cryptography — Encryption and Key Transport In order to overcome these drawbacks, Whitfield Diffie, Martin Hellman and Ralph Merkle made a revolutionary proposal based on the following idea: It is not neces- sary that the key possessed by the person who encrypts the message (that’s Alice in our example) is secret. The crucial part is that Bob, the receiver, can only decrypt using a secret key. In order to realize such a system, Bob publishes a public encryp- tion key, which is known to everyone. Bob also has a matching secret key, which is used for decryption. Thus, Bob’s key k consists of two parts, a public part, kpub, and a private one, kpr. A simple analogy for such a system is shown in Figure 6.3. This system works quite similarly to the good old mailbox on the corner of a street: Everyone can put a letter in the box, i.e., encrypt, but only a person with a private (secret) key can retrieve letters, i.e., decrypt. If we assume we have cryptosystems with such a functionality, a basic protocol for public-key encryption looks like Figure 6.4. Alice Bob kpub ←−−−−−−−−−−−− (kpub, kpr ) = k y = ekpub (x) y −−−−−−−−−−−−→ x = dkpr (y) Fig. 6.4 Basic protocol for public-key encryption By looking at that protocol one might argue that even though we can encrypt a message without a secret channel for key establishment, we still cannot exchange a key if we want to encrypt with, say, AES. However, the protocol can easily be mod- ified for this use. What we have to do is to encrypt a symmetric key, e.g., an AES key, using the public-key algorithm. Once the symmetric key has been decrypted
6.1 Symmetric vs. Asymmetric Cryptography 181 by Bob, both parties can use it to encrypt and decrypt messages using symmetric ciphers. Figure 6.5 shows a basic key transport protocol where we use AES as the symmetric cipher for illustration purposes (of course, one can use any other symmet- ric algorithm in such a protocol). The main advantage of the protocol in Figure 6.5 over the one in Figure 6.4 is that the payload is encrypted with a symmetric cipher, which tends to be much faster than an asymmetric algorithm, as we will discuss later. Alice Bob kpub ←−−−−−−−−−−−− kpub, kpr choose random kAES y = ekpub (kAES) y −−−−−−−−−−−−→ kAES = dkpr (y) encrypt message x: z = AESkAES (x) z −−−−−−−−−−−−→ x = AES−1 kAES (z) Fig. 6.5 Simple key transport protocol (with AES as an example of a symmetric cipher) In practice the symmetric key is shorter than the asymmetric key and needs to be padded in order to serve as input for the asymmetric encryption. While this can be done by applying generic padding schemes, it is beneficial to consider the combination of supplying a secret key for symmetric encryption using asymmetric encryption as one joint primitive. We call such a dedicated primitive a key encap- sulation mechanism (KEM). In contrast to the plain encryption functions we have seen before, KEMs do not expect a message as input. Instead, KEMs are capable of generationg a random value (that we use later as a symmetric secret key) on their own which is then fed into an associated asymmetric encryption function. This way, Alice just needs to invoke a single function encapsulate to create random key mate- rial which is directly asymmetrically encrypted. After Bob receives this encrypted key material, Bob runs the corresponding function decapsulate to decrypt and un- wrap the key material, which Alice and Bob then share for subsequent symmetric bulk data encryption. We will explain how KEMs are used and instantiated in detail in the context of RSA (cf. Section 7.8) and post-quantum cryptography (cf. Chap- ter 12). From the discussion so far, it has become obvious that asymmetric cryptography is a desirable tool for building security solutions. The not-so-small question remain- ing is how one can build public-key algorithms. In Chaps. 7, 8 and 9 we introduce the asymmetric schemes that are currently in use. They are all built from one com- mon principle, a one-way function. A one-way function is basically a function that
182 6 Introduction to Public-Key Cryptography is easy to compute on every input, but hard to reverse given any output of a random input. The informal definition is as follows. Definition 6.1.1 One-way function A function f is a one-way function if: 1. y = f (x) is computationally easy, and 2. x = f −1(y) is computationally infeasible. Obviously, the adjectives “easy” and “infeasible” are not particularly exact. In the sense of computational complexity theory, a function is easy to compute if it can be evaluated in polynomial time, i.e., its running time is a polynomial expression. In order to be useful in practical cryptographic schemes, the computation y = f (x) should be sufficiently fast that it does not lead to unacceptably slow execution times in an application. The computation of the inverse x = f −1(y) should be computa- tionally very costly in the average case. This means that it should be infeasible to evaluate it in any reasonable time period, say, 10,000 years, when using the best known algorithm and all computer resources on planet Earth. There are several popular one-way functions that are used by the public-key schemes currently in use. The first is the integer factorization problem, on which RSA is based. Given two large primes, it is easy to compute the product. However, it is very difficult to factor the resulting product on common computing platforms. In fact, if each of the primes has 300 or more decimal digits, the resulting product cannot be factored even with thousands of (classical) computers running for many years. Another one-way function that is used widely is the discrete logarithm prob- lem, on which schemes such as the Diffie-Hellman key exchange and elliptic curves are based. This is not quite as intuitive and is introduced in Chapter 8. We remark that with the potential of future powerful quantum computers we need to reconsider the security of any of the aforementioned one-way functions. We will discuss this problem and alternatives in Chapter 12. 6.2 Practical Aspects of Public-Key Cryptography Public-key algorithms will be introduced in subsequent chapters since there is some mathematics we must study first. However, it is very interesting to look at the prin- cipal security mechanisms that are enabled by public-key cryptography, which we address in this section.
6.2 Practical Aspects of Public-Key Cryptography 183 6.2.1 Security Mechanisms As shown earlier in this chapter, public-key schemes can be used for encryption of data. It turns out that we can do many other things with public-key algorithms. The main functions that they can provide are listed below: Main Security Mechanisms of Public-Key Algorithms: Key Establishment There are protocols for establishing secret keys over an insecure channel. Examples of such protocols include the Diffie– Hellman key exchange (DHKE), the RSA key transport protocol and key encapsulation mechanisms. Non-repudiation Providing non-repudiation can be realized with digital signature algorithms, e.g., RSA, DSA or ECDSA. Integrity Digital signature algorithms can also ensure the integrity of mes- sages. Identification We can identify entities using challenge-and-response pro- tocols together with digital signatures, e.g., in applications such as smart cards for banking or for mobile phones. Encryption We can encrypt messages using algorithms such as RSA or Elgamal. We note that encryption, message integrity and identification and can also be achieved with symmetric ciphers, but they are not good at key management and non- repudiation is extremely difficult with them. It looks as though public-key schemes can provide all functions required by modern security applications. Even though this is true in principle, the major drawback in practice is that encryption of data is very computationally intensive — or more colloquially: extremely slow — with public- key algorithms. Most block and stream ciphers can encrypt about one hundred to one thousand times faster than public-key algorithms. Thus, somewhat ironically, public-key cryptography is rarely used for the oldest application of cryptography, namely the actual encryption of data. On the other hand, symmetric algorithms are poor at providing non-repudiation and key establishment functionality. In order to use the best of both worlds, most practical protocols are hybrid protocols which incorporate both symmetric and public-key algorithms. Examples include the TLS protocol that is commonly used for secure internet connections, or many end-to-end encryption protocols used by instant messengers such as WhatsApp.
184 6 Introduction to Public-Key Cryptography 6.2.2 The Remaining Problem: Authenticity of Public Keys From the discussion so far we have seen that a major advantage of asymmetric schemes is that we can distribute public keys over insecure channels, as shown in the protocols in Figures 6.4 and 6.5. However, in practice, things are a bit more tricky because we still have to ensure the authenticity of public keys. In other words: Do we really know that a certain public key belongs to a certain person? In practice, this issue is often solved with what is called certificates. Roughly speaking, certifi- cates bind a public key to a certain identity. This is a major issue in many security applications, e.g., when doing e-commerce transactions on the internet. We discuss this topic in more detail in Section 14.4.2. Another problem, which is not as fundamental, is that public-key algorithms re- quire very long keys, resulting in slow execution times. The issue of key lengths and security is discussed below. 6.2.3 Important Public-Key Algorithms In the previous chapters, we learned about some block ciphers, DES, AES and PRESENT, and a few stream ciphers. However, there exist many other symmetric algorithms. Several hundred ciphers have been proposed over the years and many are cryptographically strong, as discussed in Section 3.7. The situation is quite dif- ferent for asymmetric algorithms. There are only three major families of public- key algorithms that are currently used in practice. They can be classified based on their underlying one-way function. We note that should quantum computers become available in the future, the three algorithm families below can all be broken. At this point, one has to use what is called post-quantum cryptography (PQC), which is a (fancy) name for alternative public-key algorithms that appear to resist attacks with quantum computers. Chapter 12 will introduce PQC schemes.