The SM4 Algorithm
SM4 is China's national block cipher standard, a 32-round design that updates one 32-bit word at a time. Learn how its byte S-box, rotate-and-XOR mixing, and compact key schedule work, and where the cipher is used.
Interactive SM4 Encryption
🔐 SM4 Encryption
The SM4 Algorithm
Introduction
SM4 is the block cipher in China’s national cryptography standards. It has the same outward shape as AES: a 128-bit block and a 128-bit key. Inside, it is a different design. Instead of transforming the whole block each round, SM4 updates just one 32-bit word per round, and it does that 32 times.
SM4 belongs to the ShangMi (SM) family of Chinese standards, alongside the SM2 public-key algorithms and the SM3 hash function. It is built into Chinese wireless networking standards, it appears in TLS cipher suites, and recent processors include instructions to speed it up.
Table of Contents
- History and Standardization
- How SM4 Works
- The Round Function T
- The S-Box
- Key Schedule
- A Worked Example
- Python Implementation
- Limitations
- Security Status
- FAQ
- References
History and Standardization
SM4 was designed at the Data Assurance and Communication Security Center of the Chinese Academy of Sciences, mainly by Lü Shuwang. The cipher was used in China’s WAPI wireless security standard, and the algorithm was declassified in January 2006. In those early days it was called SMS4.
It then moved through the standards process. It was published as the industry standard GM/T 0002 in 2012 and became the national standard GB/T 32907-2016 in August 2016. In 2021 it was added to the international standard ISO/IEC 18033-3 as an amendment. The IETF also published an informational RFC, RFC 8998, that defines TLS 1.3 cipher suites using SM4.
How SM4 Works
SM4 treats the 128-bit block as four 32-bit words, X0 to X3, read as big-endian. Each round creates one new word from the previous four:
X[i+4] = X[i] ^ T(X[i+1] ^ X[i+2] ^ X[i+3] ^ rk[i])
The three newest words and a 32-bit round key rk[i] are XORed together and passed through a function T. The result is XORed into the oldest word, and the window slides forward by one position. This is called an unbalanced Feistel network, because one word changes while three stay put. Ordinary Feistel ciphers like DES update half of the block instead.
After 32 rounds the cipher outputs the last four words in reverse order: X35, X34, X33, X32. That final reversal is what makes decryption so easy. The decryption procedure is identical to encryption, except the 32 round keys are used in the opposite order.
Interactive Visualizer
The visualizer above runs this exact algorithm. Each row shows the current window of four words. A new word appears at the right after every round, and the note names the round key that was used.
The Round Function T
The function T has two stages, applied in sequence. First comes a nonlinear step called τ (tau), then a linear mixing step called L:
- τ, the substitution. Split the 32-bit input into four bytes and replace each one using the same 8-bit S-box.
- L, the mixing. Spread the bits of the result across the word with rotations and XORs.
L(B) = B ^ (B <<< 2) ^ (B <<< 10) ^ (B <<< 18) ^ (B <<< 24)
Here <<< is a left rotation. The substitution gives nonlinearity and the mixing gives diffusion. After one application of L, every output bit depends on five input bits, and after a few rounds a change in any one input bit has reached the whole block.
The S-Box
SM4 uses a single S-box that maps one byte to another byte. It is a permutation of the numbers 0 to 255, and the specification gives it as a 16 by 16 table. For example, the input EF reads row E, column F, and produces 84.
The S-box is based on the multiplicative inverse in the finite field GF(2⁸), the same core idea as the AES S-box. The affine transforms and the polynomial basis are different, but the two S-boxes are related, and one can be computed from the other efficiently. Since it is built from field inversion, it has good resistance to differential and linear attacks.
Key Schedule
The 128-bit key is read as four words MK0 to MK3. Each is XORed with one of four fixed “family key” constants:
FK = a3b1bac6, 56aa3350, 677d9197, b27022dc
Those four results become K0 to K3. The schedule then produces 32 more words using nearly the same structure as the encryption rounds:
K[i+4] = K[i] ^ T'(K[i+1] ^ K[i+2] ^ K[i+3] ^ CK[i])
rk[i] = K[i+4]
Two things differ from the encryption round. The linear step is L', which is B ^ (B <<< 13) ^ (B <<< 23). And each round uses a constant CK[i] instead of a round key. These constants do not need a table. Byte j of CK[i] is (4i + j) × 7 mod 256, which gives 00070e15 for CK0 and 1c232a31 for CK1.
A Worked Example
This is the first example from the standard GB/T 32907-2016. The key is 0123456789abcdeffedcba9876543210, and the plaintext is the very same value.
Key schedule. The key words are 01234567 89abcdef fedcba98 76543210. XORing with FK gives K0 to K3 as a292ffa1 df01febf 99a12b0f c42410cc. The first four round keys come out as rk0 = f12186f9, rk1 = 41662b61, rk2 = 5a6ab19a, and rk3 = 7ba92077. The last one, rk31, is 9124a012.
Round 1. The input to T is X1 ^ X2 ^ X3 ^ rk0 = f002c39e. The four S-box lookups turn that into 18e992b1, and the linear step L turns that into 26d99622. XORing with X0 gives the first new word, X4 = 27fad345.
Next rounds. Continuing the pattern gives X5 = a18b4cb2, X6 = 11c1e22a, and X7 = cc13e2ee.
Result. After 32 rounds the last four words are X32 = 536e4246, X33 = 86b3e94f, X34 = d206965e, and X35 = 681edf34. The reverse transformation outputs them as X35, X34, X33, X32, so the ciphertext is:
681edf34d206965e86b3e94f536e4246
All 32 round keys and all 32 round outputs match the table printed in the standard. The standard’s second example encrypts the same block one million times in a row with the same key, and the result is 595298c7c6fd271f0402f804c33d3f66.
Python Implementation
This is a complete SM4: the real S-box, the real family key, and the real 32-round key schedule and encryption. The S-box was copied from the specification by a script, and the 32 round constants are computed from their formula.
# sm4.py
#
# SM4, the Chinese national block cipher standard (GB/T 32907-2016, also
# called SMS4). A 128-bit block and a 128-bit key. 32 rounds of an
# unbalanced Feistel network: each round updates one of four 32-bit words
# using the other three. The S-box and family key are copied from the
# specification; the 32 round constants are computed from a formula.
MASK = 0xFFFFFFFF
SBOX = bytes([
0xd6, 0x90, 0xe9, 0xfe, 0xcc, 0xe1, 0x3d, 0xb7,
0x16, 0xb6, 0x14, 0xc2, 0x28, 0xfb, 0x2c, 0x05,
0x2b, 0x67, 0x9a, 0x76, 0x2a, 0xbe, 0x04, 0xc3,
0xaa, 0x44, 0x13, 0x26, 0x49, 0x86, 0x06, 0x99,
0x9c, 0x42, 0x50, 0xf4, 0x91, 0xef, 0x98, 0x7a,
0x33, 0x54, 0x0b, 0x43, 0xed, 0xcf, 0xac, 0x62,
0xe4, 0xb3, 0x1c, 0xa9, 0xc9, 0x08, 0xe8, 0x95,
0x80, 0xdf, 0x94, 0xfa, 0x75, 0x8f, 0x3f, 0xa6,
0x47, 0x07, 0xa7, 0xfc, 0xf3, 0x73, 0x17, 0xba,
0x83, 0x59, 0x3c, 0x19, 0xe6, 0x85, 0x4f, 0xa8,
0x68, 0x6b, 0x81, 0xb2, 0x71, 0x64, 0xda, 0x8b,
0xf8, 0xeb, 0x0f, 0x4b, 0x70, 0x56, 0x9d, 0x35,
0x1e, 0x24, 0x0e, 0x5e, 0x63, 0x58, 0xd1, 0xa2,
0x25, 0x22, 0x7c, 0x3b, 0x01, 0x21, 0x78, 0x87,
0xd4, 0x00, 0x46, 0x57, 0x9f, 0xd3, 0x27, 0x52,
0x4c, 0x36, 0x02, 0xe7, 0xa0, 0xc4, 0xc8, 0x9e,
0xea, 0xbf, 0x8a, 0xd2, 0x40, 0xc7, 0x38, 0xb5,
0xa3, 0xf7, 0xf2, 0xce, 0xf9, 0x61, 0x15, 0xa1,
0xe0, 0xae, 0x5d, 0xa4, 0x9b, 0x34, 0x1a, 0x55,
0xad, 0x93, 0x32, 0x30, 0xf5, 0x8c, 0xb1, 0xe3,
0x1d, 0xf6, 0xe2, 0x2e, 0x82, 0x66, 0xca, 0x60,
0xc0, 0x29, 0x23, 0xab, 0x0d, 0x53, 0x4e, 0x6f,
0xd5, 0xdb, 0x37, 0x45, 0xde, 0xfd, 0x8e, 0x2f,
0x03, 0xff, 0x6a, 0x72, 0x6d, 0x6c, 0x5b, 0x51,
0x8d, 0x1b, 0xaf, 0x92, 0xbb, 0xdd, 0xbc, 0x7f,
0x11, 0xd9, 0x5c, 0x41, 0x1f, 0x10, 0x5a, 0xd8,
0x0a, 0xc1, 0x31, 0x88, 0xa5, 0xcd, 0x7b, 0xbd,
0x2d, 0x74, 0xd0, 0x12, 0xb8, 0xe5, 0xb4, 0xb0,
0x89, 0x69, 0x97, 0x4a, 0x0c, 0x96, 0x77, 0x7e,
0x65, 0xb9, 0xf1, 0x09, 0xc5, 0x6e, 0xc6, 0x84,
0x18, 0xf0, 0x7d, 0xec, 0x3a, 0xdc, 0x4d, 0x20,
0x79, 0xee, 0x5f, 0x3e, 0xd7, 0xcb, 0x39, 0x48,
])
FK = [0xA3B1BAC6, 0x56AA3350, 0x677D9197, 0xB27022DC]
def rol(x, n):
return ((x << n) | (x >> (32 - n))) & MASK
def ck(i):
"""Round constant i: the bytes (4i + j) * 7 mod 256 for j = 0..3."""
return int.from_bytes(bytes((4 * i + j) * 7 % 256 for j in range(4)), "big")
def tau(x):
"""Four S-boxes in parallel, one per byte."""
return int.from_bytes(bytes(SBOX[b] for b in x.to_bytes(4, "big")), "big")
def lin(b):
return b ^ rol(b, 2) ^ rol(b, 10) ^ rol(b, 18) ^ rol(b, 24)
def lin_key(b):
return b ^ rol(b, 13) ^ rol(b, 23)
def round_keys(key):
k = [int.from_bytes(key[4 * i:4 * i + 4], "big") ^ FK[i] for i in range(4)]
for i in range(32):
k.append(k[i] ^ lin_key(tau(k[i + 1] ^ k[i + 2] ^ k[i + 3] ^ ck(i))))
return k[4:]
def crypt_block(rk, block):
x = [int.from_bytes(block[4 * i:4 * i + 4], "big") for i in range(4)]
for i in range(32):
x.append(x[i] ^ lin(tau(x[i + 1] ^ x[i + 2] ^ x[i + 3] ^ rk[i])))
return b"".join(v.to_bytes(4, "big") for v in reversed(x[32:]))
def encrypt_block(key, block):
return crypt_block(round_keys(key), block)
def decrypt_block(key, block):
return crypt_block(round_keys(key)[::-1], block)
if __name__ == "__main__":
key = bytes.fromhex("0123456789abcdeffedcba9876543210")
plaintext = bytes.fromhex("0123456789abcdeffedcba9876543210")
ciphertext = encrypt_block(key, plaintext)
recovered = decrypt_block(key, ciphertext)
print(f"Key: {key.hex()}")
print(f"Plaintext: {plaintext.hex()}")
print(f"Ciphertext: {ciphertext.hex()}")
print(f"Recovered: {recovered.hex()}")
Running it produces this output:
Key: 0123456789abcdeffedcba9876543210
Plaintext: 0123456789abcdeffedcba9876543210
Ciphertext: 681edf34d206965e86b3e94f536e4246
Recovered: 0123456789abcdeffedcba9876543210
I checked this code against four sources before writing it up. It reproduces every round key and round output from the standard’s first example, not just the final ciphertext. It reproduces the standard’s one-million-encryption example exactly. It matches the vectors in the Linux kernel test suite, and it agrees with OpenSSL on 60 random key and block pairs, in both directions.
For Fun: The Same Cipher in 18 Lines
This is the same spirit as the compact bonus sections elsewhere on this site. It is not for learning the algorithm from. This version squeezes the 103-line implementation above into 18 lines. The S-box takes one line, and the substitution, the two linear layers, and the key schedule are written as tight one-liners. It needs no imports and no other files.
MASK=0xFFFFFFFF
SBOX=bytes([0xd6,0x90,0xe9,0xfe,0xcc,0xe1,0x3d,0xb7,0x16,0xb6,0x14,0xc2,0x28,0xfb,0x2c,0x05,0x2b,0x67,0x9a,0x76,0x2a,0xbe,0x04,0xc3,0xaa,0x44,0x13,0x26,0x49,0x86,0x06,0x99,0x9c,0x42,0x50,0xf4,0x91,0xef,0x98,0x7a,0x33,0x54,0x0b,0x43,0xed,0xcf,0xac,0x62,0xe4,0xb3,0x1c,0xa9,0xc9,0x08,0xe8,0x95,0x80,0xdf,0x94,0xfa,0x75,0x8f,0x3f,0xa6,0x47,0x07,0xa7,0xfc,0xf3,0x73,0x17,0xba,0x83,0x59,0x3c,0x19,0xe6,0x85,0x4f,0xa8,0x68,0x6b,0x81,0xb2,0x71,0x64,0xda,0x8b,0xf8,0xeb,0x0f,0x4b,0x70,0x56,0x9d,0x35,0x1e,0x24,0x0e,0x5e,0x63,0x58,0xd1,0xa2,0x25,0x22,0x7c,0x3b,0x01,0x21,0x78,0x87,0xd4,0x00,0x46,0x57,0x9f,0xd3,0x27,0x52,0x4c,0x36,0x02,0xe7,0xa0,0xc4,0xc8,0x9e,0xea,0xbf,0x8a,0xd2,0x40,0xc7,0x38,0xb5,0xa3,0xf7,0xf2,0xce,0xf9,0x61,0x15,0xa1,0xe0,0xae,0x5d,0xa4,0x9b,0x34,0x1a,0x55,0xad,0x93,0x32,0x30,0xf5,0x8c,0xb1,0xe3,0x1d,0xf6,0xe2,0x2e,0x82,0x66,0xca,0x60,0xc0,0x29,0x23,0xab,0x0d,0x53,0x4e,0x6f,0xd5,0xdb,0x37,0x45,0xde,0xfd,0x8e,0x2f,0x03,0xff,0x6a,0x72,0x6d,0x6c,0x5b,0x51,0x8d,0x1b,0xaf,0x92,0xbb,0xdd,0xbc,0x7f,0x11,0xd9,0x5c,0x41,0x1f,0x10,0x5a,0xd8,0x0a,0xc1,0x31,0x88,0xa5,0xcd,0x7b,0xbd,0x2d,0x74,0xd0,0x12,0xb8,0xe5,0xb4,0xb0,0x89,0x69,0x97,0x4a,0x0c,0x96,0x77,0x7e,0x65,0xb9,0xf1,0x09,0xc5,0x6e,0xc6,0x84,0x18,0xf0,0x7d,0xec,0x3a,0xdc,0x4d,0x20,0x79,0xee,0x5f,0x3e,0xd7,0xcb,0x39,0x48])
FK=[0xA3B1BAC6,0x56AA3350,0x677D9197,0xB27022DC]
rol=lambda x,n: ((x<<n)|(x>>(32-n)))&MASK
ck=lambda i: int.from_bytes(bytes((4*i+j)*7%256 for j in range(4)),"big")
tau=lambda x: int.from_bytes(bytes(SBOX[b] for b in x.to_bytes(4,"big")),"big")
lin=lambda b: b^rol(b,2)^rol(b,10)^rol(b,18)^rol(b,24)
lin_key=lambda b: b^rol(b,13)^rol(b,23)
def round_keys(key):
k=[int.from_bytes(key[4*i:4*i+4],"big")^FK[i] for i in range(4)]
for i in range(32): k.append(k[i]^lin_key(tau(k[i+1]^k[i+2]^k[i+3]^ck(i))))
return k[4:]
def crypt_block(rk,block):
x=[int.from_bytes(block[4*i:4*i+4],"big") for i in range(4)]
for i in range(32): x.append(x[i]^lin(tau(x[i+1]^x[i+2]^x[i+3]^rk[i])))
return b"".join(v.to_bytes(4,"big") for v in reversed(x[32:]))
encrypt_block=lambda key,block: crypt_block(round_keys(key),block)
decrypt_block=lambda key,block: crypt_block(round_keys(key)[::-1],block)
key=bytes.fromhex("0123456789abcdeffedcba9876543210"); pt=bytes.fromhex("0123456789abcdeffedcba9876543210"); ct=encrypt_block(key,pt); rec=decrypt_block(key,ct); print(f"Key: {key.hex()}"); print(f"Plaintext: {pt.hex()}"); print(f"Ciphertext: {ct.hex()}"); print(f"Recovered: {rec.hex()}")
Running it prints the same four lines as the readable version, including the standard test vector ciphertext:
Key: 0123456789abcdeffedcba9876543210
Plaintext: 0123456789abcdeffedcba9876543210
Ciphertext: 681edf34d206965e86b3e94f536e4246
Recovered: 0123456789abcdeffedcba9876543210
I checked it against the readable code on 500 random keys and blocks. Encryption matched every time, and decryption recovered every block. Chaining 1,000 encryptions of the test block also gave the same result in both versions. The S-box and the FK constants are identical.
Limitations
This is a faithful teaching version of the cipher, not a production one:
- Single 128-bit block only. There is no mode of operation for longer messages and no padding for partial blocks.
- Table lookups depend on secret data. The S-box lookup uses secret values as indexes, so real software must consider cache-timing leaks that this plain code does not address.
- Slow by design. Hardware instructions and bitsliced code run SM4 far faster than this readable loop.
- No side-channel hardening. There is no protection against power or fault attacks.
Security Status
No practical attack on the full 32 rounds is known. The best published cryptanalysis, using linear and differential techniques, reaches about 23 of the 32 rounds, and some summaries quote 22. That leaves a comfortable margin. SM4 is generally regarded as secure when used correctly in a proper mode of operation.
In practice, the choice between SM4 and AES is usually about regulations and compatibility rather than security. Where Chinese standards or products require SM4, it works as a strong 128-bit cipher. Elsewhere, AES is the usual default because of its wider hardware support. SM4 has its own, though. It is part of the ARMv8.4-A extensions, and RISC-V standardized an SM4 extension called Zksed in 2021.
FAQ
Is SM4 secure?
As far as anyone has published, yes. The best attacks reach 22 of 32 rounds, and no practical attack on the full cipher exists. It has a 128-bit block, so it avoids the birthday-bound limits that affect 64-bit ciphers such as DES and Blowfish.
How is SM4 different from AES?
Both use a 128-bit block and a 128-bit key and an S-box built from field inversion. AES is a substitution-permutation network that transforms the whole block each round, with 10 rounds. SM4 is an unbalanced Feistel network that updates one word per round, with 32 rounds.
Why does decryption just reverse the round keys?
The final output step reverses the order of the last four words. That makes the whole procedure self-inverse in structure, so running the same rounds with the keys in reverse order undoes encryption.
Where is SM4 used?
It is used in the Chinese WAPI wireless standard and in TLS cipher suites defined by RFC 8998. It is available in OpenSSL and the Linux kernel, and ARMv8.4-A and the RISC-V Zksed extension include instructions for it.
What do the SM2, SM3, and SM4 names mean?
They are the public-key, hash, and block cipher members of the ShangMi family of Chinese cryptographic standards. SM4 is the symmetric cipher in that set.
References
-
“SM4 block cipher algorithm.” Chinese national standard GB/T 32907-2016, 2016.
-
Tse, R., Wong, W., and Saarinen, M-J. “The SM4 Blockcipher Algorithm And Its Modes Of Operations.” Internet-Draft draft-ribose-cfrg-sm4, 2018. Available at: https://datatracker.ietf.org/doc/draft-ribose-cfrg-sm4/
-
Yang, P. “ShangMi (SM) Cipher Suites for TLS 1.3.” RFC 8998, 2021. Available at: https://www.rfc-editor.org/rfc/rfc8998
-
Wikipedia. “SM4 (cipher).” Available at: https://en.wikipedia.org/wiki/SM4_(cipher)