Skip to content
clintvalPublic

About

Streaming BGZF compression with on-the-fly tabix and CSI indexing

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

pybgzf

Build Status Python Versions License: MIT Language

Streaming BGZF compression with on-the-fly tabix and CSI indexing.

Install with pip or uv:

pip install pybgzf

Introduction

pybgzf writes blocked gzip (BGZF) files from Python and builds their tabix (.tbi) or CSI (.csi) index while it writes, so there is no second pass with tabix. Its core is Rust: blocks are compressed with the multithreaded, libdeflate-backed bgzf crate and indexes are written with noodles. It needs neither pysam nor htslib.

Quick Start

Write sorted BED lines as text and the index appears next to the file on close:

>>> from pathlib import Path
>>> from tempfile import mkdtemp
>>>
>>> import pybgzf
>>> from pybgzf import Columns, IndexFormat
>>>
>>> directory = Path(mkdtemp())
>>> path = directory / "features.bed.gz"
>>>
>>> with pybgzf.writer(path, index=IndexFormat.TBI, columns=Columns.BED) as handle:
...     _ = handle.write("chr1\t100\t200\tgene-a\n")
...     _ = handle.write("chr1\t150\t300\tgene-b\n")
...     _ = handle.write("chr2\t10\t20\tgene-c\n")
>>>
>>> sorted(file.name for file in directory.iterdir())
['features.bed.gz', 'features.bed.gz.tbi']

The output is ordinary BGZF, so gzip -d, bgzip -d, and tabix features.bed.gz chr1:120-160 all read it. pybgzf.BGZF_SUFFIXES holds the file name suffixes of BGZF files, for deciding when to write one:

>>> pybgzf.BGZF_SUFFIXES
('.gz', '.bgz', '.bgzf')
>>> path.name.endswith(pybgzf.BGZF_SUFFIXES)
True

Stream bytes to a pipe or any binary file object, with columns inferred from the content:

>>> import io
>>>
>>> stream = io.BytesIO()
>>> with pybgzf.BgzfWriter(stream, threads=4, index=IndexFormat.CSI, index_path=directory / "calls.csi", columns=pybgzf.INFER) as writer:
...     _ = writer.write(b"##fileformat=VCFv4.3\n")
...     writer.columns == Columns.VCF
True

Read the BED file back on several threads, or query a region with 0-based, half-open coordinates:

>>> with pybgzf.reader(path, threads=4) as handle:
...     handle.readline()
'chr1\t100\t200\tgene-a\n'

>>> with pybgzf.IndexedReader(path) as reader:
...     list(reader.query("chr1", 180, 190))
['chr1\t100\t200\tgene-a', 'chr1\t150\t300\tgene-b']

Development and Testing

See the contributing guide for more information.

The multithreaded, position-tracking writer is adapted from fgumi, and indexing follows htslib; see NOTICE.

About

Streaming BGZF compression with on-the-fly tabix and CSI indexing

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Contributors

Languages