![Introducing Enhanced Alert Actions and Triage Functionality](https://cdn.sanity.io/images/cgdhsj6q/production/fe71306d515f85de6139b46745ea7180362324f0-2530x946.png?w=800&fit=max&auto=format)
Product
Introducing Enhanced Alert Actions and Triage Functionality
Socket now supports four distinct alert actions instead of the previous two, and alert triaging allows users to override the actions taken for all individual alerts.
html-encoding-sniffer
Advanced tools
Package description
The html-encoding-sniffer npm package is designed to determine the encoding of HTML documents. It does this by examining the byte stream of the document, looking for any encoding declarations in the form of a meta tag or an HTTP header. This is particularly useful for applications that need to correctly interpret or display HTML content from various sources, ensuring that text is properly encoded and displayed.
Sniffing HTML encoding from HTTP headers
This feature allows you to determine the encoding of an HTML document by examining the HTTP headers. The 'transportLayerEncodingLabel' option is used to specify the encoding declared in the HTTP headers.
"use strict";
const htmlEncodingSniffer = require('html-encoding-sniffer');
const encoding = htmlEncodingSniffer(byteStream, { transportLayerEncodingLabel: 'utf-8' });
Sniffing HTML encoding from a meta tag
This feature enables the detection of the document's encoding by looking for a meta tag within the HTML that specifies the encoding. The 'defaultEncoding' option allows you to specify a fallback encoding in case no encoding is declared in the document.
"use strict";
const htmlEncodingSniffer = require('html-encoding-sniffer');
const encoding = htmlEncodingSniffer(byteStream, { defaultEncoding: 'windows-1252' });
iconv-lite is a package that provides encoding and decoding of text in various character sets. Unlike html-encoding-sniffer, which is specifically designed for sniffing HTML document encodings, iconv-lite supports a broader range of encodings and can be used for general text conversion purposes.
jschardet is a character encoding detector, similar to the functionality provided by html-encoding-sniffer. However, jschardet is based on the universalchardet library and can be used to detect the encoding of any text, not just HTML documents. It offers a more general approach to encoding detection.
Readme
This package implements the HTML Standard's encoding sniffing algorithm in all its glory. The most interesting part of this is how it pre-scans the first 1024 bytes in order to search for certain <meta charset>
-related patterns.
const htmlEncodingSniffer = require("html-encoding-sniffer");
const fs = require("fs");
const htmlBuffer = fs.readFileSync("./html-page.html");
const sniffedEncoding = htmlEncodingSniffer(htmlBuffer);
The returned value will be a canonical encoding name (not a label). You might then combine this with the whatwg-encoding package to decode the result:
const whatwgEncoding = require("whatwg-encoding");
const htmlString = whatwgEncoding.decode(htmlBuffer, sniffedEncoding);
You can pass two potential options to htmlEncodingSniffer
:
const sniffedEncoding = htmlEncodingSniffer(htmlBuffer, {
transportLayerEncodingLabel,
defaultEncoding
});
These represent two possible inputs into the encoding sniffing algorithm:
transportLayerEncodingLabel
is an encoding label that is obtained from the "transport layer" (probably a HTTP Content-Type
header), which overrides everything but a BOM.defaultEncoding
is the ultimate fallback encoding used if no valid encoding is supplied by the transport layer, and no encoding is sniffed from the bytes. It defaults to "windows-1252"
, as recommended by the algorithm's table of suggested defaults for "All other locales" (including the en
locale).This package was originally based on the excellent work of @nicolashenry, in jsdom. It has since been pulled out into this separate package.
FAQs
Sniff the encoding from a HTML byte stream
The npm package html-encoding-sniffer receives a total of 20,196,801 weekly downloads. As such, html-encoding-sniffer popularity was classified as popular.
We found that html-encoding-sniffer demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 6 open source maintainers collaborating on the project.
Did you know?
Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.
Product
Socket now supports four distinct alert actions instead of the previous two, and alert triaging allows users to override the actions taken for all individual alerts.
Security News
Polyfill.io has been serving malware for months via its CDN, after the project's open source maintainer sold the service to a company based in China.
Security News
OpenSSF is warning open source maintainers to stay vigilant against reputation farming on GitHub, where users artificially inflate their status by manipulating interactions on closed issues and PRs.