Text Splitter
Splits text into chunks that fit a limit, counts the SMS segments properly, and reserves room for the numbering before it splits rather than after.
Chunks 2
Longest 152 characters
Limit 153 for SMS, one part of a longer message
Room used by numbering 6 characters a chunk
1 152 characters
Our winter sale starts on Friday and runs for ten days. Everything in the outlet is reduced, including the ranges we do not normally discount, and (1/2)
2 88 characters
members get an extra hour before it opens to everyone else. Reply STOP to opt out. (2/2)
As an SMS, before splitting
encoding GSM-7
units 229 of 153
parts 2
An SMS is 1,120 bits, not 160 characters. It is 160 when every character
is in the GSM 03.38 alphabet at seven bits each, and 70 when any single
one is not, because one character outside that alphabet moves the whole
message to UCS-2 at sixteen bits.
This text is entirely GSM-7, so it gets the full 160 units. Adding one
emoji, one curly quote or one en dash drops the limit to 70 and can turn
one message into three.
A long message is split by the sender and rejoined by the handset, and
each part carries a six-byte header saying which part it is. That header
costs seven GSM-7 characters or three UCS-2 characters from every part,
so the working limits are 153 and 67 rather than 160 and 70.
Twelve characters cost two units each: the square and curly brackets,
the backslash, the caret, the tilde, the vertical bar and the euro sign.
They are reached through an escape, which is why a message full of
brackets runs short of room sooner than the count suggests.
Numbering is inside the limit, not outside it. " (10/10)" is eight
characters taken off every chunk, which is why it is reserved before the
split here rather than added afterwards.
Counting is by code point, so a chunk never ends between the halves of a
surrogate pair. Split by UTF-16 unit and an emoji at the boundary
arrives as two replacement characters, which is the most common way a
split message looks broken.
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
An SMS is not 160 characters. It is 1,120 bits.
It is 160 characters when every one of them is in the GSM 03.38 alphabet, packed seven bits each. One character outside that alphabet switches the whole message to UCS-2 at sixteen bits, and the limit drops to 70. That is why adding a single emoji to a 158-character message turns it into three messages rather than one, and why a curly quote pasted from a word processor can double what a campaign costs.
So this splits text to a limit and tells you which limit actually applies.
How to use
- Paste the text.
- Pick a limit, or set your own. SMS is 153 rather than 160 for a reason explained below.
- Turn numbering on if the chunks will be sent separately. Its length is taken off the limit before the split, not after.
Example
A 229-character promotional message at the SMS part limit:
Chunks 2
Longest 152 characters
Limit 153 for SMS, one part of a longer message
Room used by numbering 6 characters a chunk
1 152 characters
Our winter sale starts on Friday and runs for ten days. Everything in the
outlet is reduced, including the ranges we do not normally discount, and (1/2)
2 88 characters
members get an extra hour before it opens to everyone else. Reply STOP to
opt out. (2/2)
Before splitting, that text is GSM-7: 229 units, which is two parts. Replace one full stop in it with an en dash and the same text reads:
As an SMS, before splitting
encoding UCS-2
units 230 of 67
parts 4
outside GSM-7 —
One character longer, and twice the messages.
Pitfalls
153, not 160. A message too long for one part is split by the sender and rejoined by the handset, and each part carries a six-byte header saying which part it is and how many there are. That header costs seven GSM-7 characters or three UCS-2 characters out of every part, so the working limits for a multipart message are 153 and 67.
One character decides the encoding for all of them. There is no mixing. A single emoji, en dash, curly quote, ellipsis character or non-Latin letter moves the entire message to UCS-2, so the fix is to find that one character rather than to shorten the text.
Twelve GSM-7 characters cost two units each. The square and curly brackets, the backslash, the caret, the tilde, the vertical bar and the euro sign live in an extension table reached through an escape. A message full of brackets runs out of room sooner than its character count suggests.
Numbering is inside the limit. ” (10/10)” is eight characters taken off every chunk. Adding it after the split is the most common bug in a splitter, and it produces chunks that are one or two characters over.
Count by code point, not by UTF-16 unit. An emoji outside the basic plane is two UTF-16 units, so slicing a JavaScript string at a fixed index can cut it in half and produce two replacement characters. That is why a split message sometimes arrives with a black diamond in the middle of it.
An SMS is not the only limit that lies. A post on X weights URLs at a fixed length and counts some scripts differently, and a meta description is truncated by pixel width rather than by character count. The presets here are the conventional numbers, not guarantees about someone else’s renderer.
Compatibility
Everything runs in the browser: nothing is uploaded and nothing is stored. That matters for a message list, which is customer data.
The GSM 03.38 basic alphabet is the full 127 printable characters plus the escape, and the extension table is the nine characters reached through it. Both are checked against the published table in the test suite, including the euro sign, which is the extended character people meet first.
Word-boundary splitting keeps whole words together and cuts a word longer than the limit rather than allowing an overflow, because a URL or a hash has to go somewhere. Counting is by code point, checked with an emoji sitting exactly on a boundary.
For a WordPress form, the same counting rules apply to anything that ends up at an SMS gateway. Most gateways bill per part, so the parts count here is the number to watch rather than the character count.