Line Counter
Counts lines, blanks, duplicates and the longest one, and explains why the answer disagrees with wc -l when the file has no trailing newline.
Lines 5
`wc -l` would say 4
why the last line has no newline after it, and `wc -l` counts newlines
With something on them 4
Empty 1
Longest line 21 characters, line 2
Indented with tabs 1
Line endings
CRLF, Windows 0
LF, Unix 4
CR alone, classic Mac 0
mostly LF
ends with a newline no ← POSIX says it should
This text has 5 lines and 4 newline characters, because the last line
has no newline after it. POSIX defines a line as characters ending with
a newline, so strictly the last one is an incomplete line, `wc -l`
reports 4, and git shows "\ No newline at end of file" in the diff.
The missing final newline is not pedantry. Concatenating two such files
joins the last line of the first to the first line of the second, a
shell loop reading line by line can drop the last line entirely, and
every commit that touches the end of the file shows a spurious change to
that line.
One ending throughout, which is what a `.gitattributes` with `text=auto`
and an editorconfig are for. Mixed endings inside one file are what
produce a diff where every line changed and nothing looks different.
Lines that look empty and are not are a common source of a wrong count:
a line holding two spaces is not blank to anything counting characters,
and it is blank to a person reading it. Both numbers are above.
`wc -l` counts newlines, `grep -c ""` counts lines the way a person
does, and `awk "END{print NR}"` agrees with grep. If a script and a
person disagree about how long a file is, this is almost always why.
A character count is not a byte count. Non-ASCII characters take more
than one byte in UTF-8, so `wc -c` and the longest-line figure here will
differ on any file with an accent or an emoji in it; `wc -m` is the one
that counts characters.
Output is valid and updates as you type.
Fix the highlighted fields to update the output.
If a script says your file has 4 lines and you can see 5, neither of you is wrong. wc -l counts
newline characters, and a file whose last line has no newline after it has one fewer newline than
it has lines.
That is not a bug in wc. POSIX defines a line as a sequence of characters ending with a newline,
so a file that does not end with one ends with an incomplete line. It is also why git prints
\ No newline at end of file in a diff, and why every commit that touches the end of such a file
shows a change to a line that did not change.
So this counts both, and says which is which.
How to use
- Paste the text or the file contents exactly, including the trailing newline or its absence.
- Read the two counts and the reason they differ.
- Check the line endings if the file has been through more than one machine.
Example
Five lines of JavaScript saved on Windows, with no newline at the end:
Lines 5
`wc -l` would say 4
why the last line has no newline after it, and `wc -l` counts newlines
With something on them 4
Empty 1
Longest line 21 characters, line 2
Indented with tabs 1
Line endings
CRLF, Windows 4
LF, Unix 0
CR alone, classic Mac 0
mostly CRLF
ends with a newline no ← POSIX says it should
Pitfalls
wc -l and grep -c "" disagree, on purpose. wc -l counts newlines. grep -c "" counts
lines the way a person does, and so does awk 'END{print NR}'. When a script and a person disagree
about the length of a file, this is almost always why.
The missing final newline is not pedantry. Concatenating two such files joins the last line of
the first to the first line of the second. A shell loop using while read drops the last line
entirely unless it is written carefully. And the diff noise is permanent until somebody fixes the
file.
Mixed line endings mean two machines. A file with both CRLF and LF has been edited on Windows
and on something else. A naive split( '\n' ) then leaves a stray carriage return at the end of
every Windows line, and "hello\r" !== "hello", so the same line compares unequal, sorts
differently, and hashes differently.
A bare CR is still out there. Mac OS before 2001 used it alone. It shows up in old exports and in files written by very old software, and it is the ending most likely to make a tool report the whole file as one line.
Lines that look empty are not always empty. A line holding two spaces is blank to a reader and not to anything counting characters. Both counts are above, because a deduplication or a “remove blank lines” step that only checks for an empty string leaves them behind.
Characters are not bytes. wc -c counts bytes and wc -m counts characters. Any accent or
emoji makes them differ, and the longest-line figure here is characters.
A trailing newline is not an empty last line. Splitting "a\nb\n" on the newline gives three
elements, the last of them empty, and there are two lines. Counting the elements is the most common
off-by-one in a line counter.
Compatibility
Everything runs in the browser: nothing is uploaded and nothing is stored.
All three endings are recognised and counted separately, and a CRLF is counted once rather than as a CR plus an LF. The longest line is measured in characters by code point, so an emoji counts as one rather than as two UTF-16 units.
The test suite covers the disagreement with wc -l in both directions, mixed endings, whitespace-only
lines, a file that is a single line with no break at all, and an empty string.
For a repository, a .gitattributes with * text=auto normalises endings on commit, and an
.editorconfig with insert_final_newline = true and end_of_line = lf stops the problem
recurring. Both are worth adding once rather than fixing files repeatedly.
Frequently asked questions
How do I add the missing final newline?
echo >> file appends one.Which line ending should I use?
.gitattributes handle the checkout.Why does my file show as one line?
Does the count include the blank line at the end?
Is there a limit on size?
grep -c "" and read the note above about why that disagrees with wc -l.