What is chardet?
The chardet npm package is a character encoding detector library, which allows you to determine the encoding of a given piece of text or a file. It is based on the character detection component of the ICU (International Components for Unicode) project and can be useful when dealing with text data that does not have encoding information.
What are chardet's main functionalities?
Detecting encoding of a text buffer
This code reads a file and uses chardet to detect the encoding of its content. The 'detect' function takes a buffer and returns the name of the encoding it believes the text is in.
const chardet = require('chardet');
const fs = require('fs');
fs.readFile('/path/to/file', (err, data) => {
if (err) throw err;
const encoding = chardet.detect(data);
console.log(encoding);
});
Detecting encoding with confidence
This code creates a buffer from a string and uses chardet's 'detectAll' function to get an array of possible encodings along with their confidence scores.
const chardet = require('chardet');
const buffer = Buffer.from('Some text with unknown encoding');
const result = chardet.detectAll(buffer);
console.log(result);
Detecting encoding of a file stream
This code creates a read stream from a file and uses chardet's 'detectStream' function to detect the encoding of the streamed content asynchronously.
const chardet = require('chardet');
const fs = require('fs');
const stream = fs.createReadStream('/path/to/file');
chardet.detectStream(stream).then(encoding => {
console.log(encoding);
});
Other packages similar to chardet
iconv-lite
iconv-lite is a character encoding conversion library. Unlike chardet, which detects the encoding, iconv-lite is used to convert from one encoding to another. It supports many encodings and is often used in conjunction with chardet to first detect the encoding and then convert the text.
jschardet
jschardet is a port of the python library chardet. It serves the same purpose as the chardet npm package, which is to detect the character encoding of text. The main difference may be in the implementation details and the specific encodings supported by each library.
encoding
The encoding npm package is another library for encoding and decoding text. It provides a simpler API for converting between encodings but does not have the detection capabilities of chardet. It's often used when the encoding is already known.
chardet
Chardet is a character detection module for NodeJS written in pure Javascript.
Module is based on ICU project http://site.icu-project.org/, which uses character
occurency analysis to determine the most probable encoding.
Installation
npm i chardet
Usage
To return the encoding with the highest confidence:
var chardet = require('chardet');
chardet.detect(Buffer.from('hello there!'));
chardet.detectFile('/path/to/file', function(err, encoding) {});
chardet.detectFileSync('/path/to/file');
To return the full list of possible encodings:
var chardet = require('chardet');
chardet.detectAll(Buffer.from('hello there!'));
chardet.detectFileAll('/path/to/file', function(err, encoding) {});
chardet.detectFileAllSync('/path/to/file');
Working with large data sets
Sometimes, when data set is huge and you want to optimize performace (in tradeoff of less accuracy),
you can sample only first N bytes of the buffer:
chardet.detectFile('/path/to/file', { sampleSize: 32 }, function(err, encoding) {});
Supported Encodings:
- UTF-8
- UTF-16 LE
- UTF-16 BE
- UTF-32 LE
- UTF-32 BE
- ISO-2022-JP
- ISO-2022-KR
- ISO-2022-CN
- Shift-JIS
- Big5
- EUC-JP
- EUC-KR
- GB18030
- ISO-8859-1
- ISO-8859-2
- ISO-8859-5
- ISO-8859-6
- ISO-8859-7
- ISO-8859-8
- ISO-8859-9
- windows-1250
- windows-1251
- windows-1252
- windows-1253
- windows-1254
- windows-1255
- windows-1256
- KOI8-R
Currently only these encodings are supported, more will be added soon.