Detect programming languages from a filename, source content, or both. The Go library runs offline and supports WebAssembly. Results include candidate languages, confidence, and conflicting evidence.
go get github.com/git-pkgs/languagesPass a filename and source bytes to Detect:
package main
import (
"fmt"
"github.com/git-pkgs/languages"
)
func main() {
source := []byte("const count = 1;\n")
result := languages.Detect("app.ts", source)
fmt.Println(result.Language, result.Confidence)
// Output: TypeScript low
}Use an empty filename for content-only detection, or nil content to use the
filename alone:
source := []byte("use strict;\nuse warnings;\n")
fmt.Println(languages.Detect("", source).Language) // Perl
fmt.Println(languages.Detect("main.go", nil).Language) // GoDetect performs no I/O and examines all supplied bytes. You can pass a prefix
when the rest of the file is unavailable. Only the final path component is used;
conventional filenames are case-sensitive and extensions are case-insensitive.
The exceptions are .C, which selects C++, and .H, which has C-family candidates.
For a file, use AnalyzeReader to read to EOF with bounded memory:
func detectFile(ctx context.Context, name string) (languages.Result, error) {
file, err := os.Open(name)
if err != nil {
return languages.Result{}, err
}
defer file.Close()
var content languages.Analysis
options := languages.ReadOptions{Filename: name}
if err := languages.AnalyzeReader(ctx, file, options, &content); err != nil {
return languages.Result{}, err
}
return content.Detect(name), nil
}Set ReadOptions{Bytes: 1024} to read at most 1 KiB; the default reads to EOF.
The reader stops exactly at the budget, so reaching it before EOF leaves
content.Prefix true. Use ReadOptions{Prefix: true} if the reader contains
truncated input.
Read errors and cancellation clear the result; cancellation is checked between
reads, and a reader blocked inside Read continues until it returns.
Set ReadOptions.Filename when the result will be detected under one filename.
This limits extension heuristics to that name without changing content evidence.
Call content.Detect(name) to combine the analysis with the filename. Leave
Filename empty to retain extension heuristics for reuse under different names.
Language is languages.Unknown when the result is ambiguous or empty. Check
Candidates to distinguish an ambiguous result from an empty one:
result := languages.Detect("source.pl", nil)
switch {
case result.Conflict:
fmt.Println("conflicting language evidence")
case result.Candidates.Empty():
fmt.Println("no language detected")
case result.Language == languages.Unknown:
names, err := json.Marshal(result.Candidates)
if err != nil {
panic(err)
}
fmt.Printf("ambiguous: %s\n", names)
default:
fmt.Println(result.Language)
}The example prints ambiguous: ["Perl","Raku","Prolog"]. Content such as
use strict; can resolve the .pl ambiguity to Perl. Check a candidate with
result.Candidates.Has(languages.Perl), or count them with
result.Candidates.Len().
Language.String() returns the name, and languages.Parse(name) accepts
canonical names and aliases regardless of case, including common-lisp.
Conflict indicates disagreement between strong content evidence and the
filename, or between a declaration and strong syntax evidence. For example, a
Ruby shebang in script.py produces a conflict. Shared syntax, such as JavaScript
that is also valid TypeScript, can be narrowed by the filename.
Confidence describes the evidence supporting the result:
none: content and filename both unmatched.low: a filename, weak syntax, statistical selection, or conflicting evidence.medium: stronger syntax evidence.high: a shebang or editor modeline, or multiple rules including a strong signal.
For a confidence threshold, compare
result.Confidence.Rank() >= languages.Medium.Rank(). High confidence can
have several candidates, so check Language as well when you need a single
selection. Confidence is an ordinal rank rather than a probability.
Statistical is true when the selection comes from the token classifier;
these results have low confidence.
Use Analyze when you have the complete content, need matched rules, or want to
detect the same content under several filenames:
source := []byte("use strict;\nuse warnings;\n")
var content languages.Analysis
languages.Analyze(source, true, &content)
fmt.Println(content.Result().Language) // Perl
fmt.Println(content.Detect("source.pl").Language) // Perl
fmt.Println(content.Detect("source.py").Conflict) // true
for _, match := range content.Signals[:content.Count] {
evidence := match.Evidence()
fmt.Println(evidence.ID, evidence.Description, evidence.Offset)
}Pass true only when the supplied bytes are the complete file. Use false for
a prefix with an unavailable suffix. Detect always treats content as a prefix.
content.Prefix records whether the input is incomplete; content.Bytes is
the number of original bytes examined.
Emacs and Vim modelines are checked in the first five lines. The last five lines
are checked when Analyze or AnalyzeReader receives the complete file.
Analyze resets its destination on each call and does not retain the input
buffer. Reuse the buffer and analysis for successive files, with a separate
destination for each concurrent call. Result and Detect read the analysis
without changing it or rescanning the source.
Rule evidence contains an ID, description, candidate languages, and a byte
offset into the original input. Statistical scores lack source offsets and are
omitted from Signals. Keep cached analyses in memory; JSON encoding drops
their classifier state.
go install github.com/git-pkgs/languages/cmd/languages@latestFor a file, the CLI writes JSON and combines the filename with
content by default. Pass a file, or pipe content to stdin and supply its name
with -name:
languages source.pl
printf 'use strict;\n' | languages -name source.plUse -mode content to ignore the filename, or -mode path to inspect a filename
without opening the file:
languages -mode content source.pl
languages -mode path -name source.plThe path-only example returns:
{
"confidence": "low",
"candidates": ["Perl", "Raku", "Prolog"],
"bytes_examined": 0,
"prefix": false,
"path_evidence": {
"reason": "extension",
"candidates": ["Perl", "Raku", "Prolog"]
}
}language is omitted for ambiguous and empty results. Unknown, ambiguous,
and conflicting results exit successfully, so scripts must inspect the JSON.
Argument, file-reading, and output errors produce a nonzero exit status.
The CLI reads full files and stdin to EOF by default. Set a read budget with
-bytes, or use -bytes 0 for full input. This command reads up to 1 KiB:
languages -bytes 1024 source.plReaching the budget without EOF sets prefix: true. For input that was truncated
before reaching stdin, use -prefix:
languages -prefix -name source.pl < exported-prefixPass a directory to see language totals for the whole tree and each subdirectory:
languages ./projectFor a project with Go and TypeScript files, output looks like:
project/ Go 75.0%, TypeScript 25.0% (3 files, 4.0 KiB)
├── api/ Go 100.0% (2 files, 3.0 KiB)
│ └── lib/ Go 100.0% (1 file, 1.0 KiB)
└── web/ TypeScript 100.0% (1 file, 1.0 KiB)
Each directory includes all descendant files, and percentages use full file
sizes even when detection reads only a prefix. Unknown, ambiguous, conflicting, and
binary files have separate totals and remain in the percentage denominator.
partial counts files whose content exceeds the read limit.
Scan a subdirectory on its own, limit the displayed depth, or request JSON:
languages ./project/api
languages -depth 1 ./project
languages -json ./project
languages -bytes 1024 ./project-depth 0 shows only the root; the default shows every directory containing
included files. Depth limits affect text and JSON output without changing the
totals. JSON contains path, summary, and children; each summary contains
file counts and byte totals, with language names under languages. JSON sizes
are integer byte counts so callers can calculate their own shares.
Scans work without Git metadata. They skip .git, symlinks, and non-regular
files, and include untracked, ignored, vendored, and generated files. Directory
scans use file sizes to identify complete input, including files exactly as
large as the read limit.
In Go, Scan accepts an fs.FS:
root, err := os.OpenRoot("project")
if err != nil {
panic(err)
}
defer root.Close()
tree, err := languages.Scan(context.Background(), root.FS(), languages.ScanOptions{})
if err != nil {
panic(err)
}
fmt.Println(tree.Root().Summary)
if api, ok := tree.Subtree("api"); ok {
fmt.Println(api.Summary)
}Use os.OpenRoot("project/api") or fs.Sub(root.FS(), "api") to scan only a
subdirectory. For a mutable or untrusted directory, prefer os.Root; see the
Scan documentation for confinement caveats.
Set ScanOptions.Bytes for a read budget, and supply ScanOptions.Exclude to
omit files or directories. It receives root-relative paths; returning true for
a directory skips its contents:
options := languages.ScanOptions{
Exclude: func(name string, entry fs.DirEntry) bool {
return entry.IsDir() && entry.Name() == "vendor"
},
}Pass these options as the third argument to Scan; if a file read fails or
the context is cancelled, the scan returns an error and discards the partial
tree.
If your application already traverses files, feed its analyses into a Tree
instead of scanning again:
var tree languages.Tree
source := []byte("package main\n")
var analysis languages.Analysis
languages.Analyze(source, true, &analysis)
if err := tree.Add("api/main.go", int64(len(source)), &analysis); err != nil {
panic(err)
}
fmt.Println(tree.Root().Summary)For prefix analyses, pass the full file size to Add. Paths must be unique,
slash-separated, and relative to the tree root, without . or .. components.
You can reuse the content buffer and analysis after each call. Root and
Subtree return independent snapshots with children sorted by path and languages
sorted by size, then file count and name. Subtree paths remain relative to the
original root; empty directories are omitted.
Detection matches byte patterns and token frequencies without validating syntax. Comment and string handling is partial, including for heredocs. Embedded languages, minified code, and short fragments may be missed or misidentified.
Go template actions such as {{.Title}} are detected in text and markup;
custom delimiters are ignored.
Language metadata and extension heuristics come from Linguist, as do the
classifier training samples. Extension heuristics inspect the first 50 KiB;
syntax rules and the token classifier process all supplied content. Binary
content sets Analysis.Binary and leaves the language unselected.
BOM-marked UTF-16 and UTF-32 are decoded before language analysis. Evidence offsets refer to the original bytes, and a prefix may end inside a code point. Malformed BOM-marked input is rejected. Other non-UTF-8 input is analyzed as raw bytes when its encoding is unrecognized; detection quality may be lower.
MIT. Imported Linguist metadata, heuristics, and sample-derived model data are covered by the upstream MIT notice.