Skip to content

Add support for HEIF based images - #2633

Open
ynse01 wants to merge 984 commits into
SixLabors:mainfrom
ynse01:heic-support
Open

ynse01 wants to merge 984 commits into
SixLabors:mainfrom
ynse01:heic-support

Conversation

@ynse01

@ynse01 ynse01 commented Dec 29, 2023 •

Copy link
Copy Markdown
Contributor

Prerequisites

  • I have written a descriptive pull-request title
  • I have verified that there are no overlapping pull-requests open
  • I have verified that I am following the existing coding patterns and practice as demonstrated in the repository. These follow strict Stylecop rules 👮.
  • I have provided test coverage for my change (where applicable)

Description

Adds managed AVIF support: a HEIF container codec with AV1 compression. Fixes #1320.

  • AVIF only. AV1 is the only supported compression. HEVC (HEIC) is not included, because HEVC is not royalty-free.
  • libaom and libavif based. The AV1 encoder is a port of libaom, and its default settings follow libavif. The libaom license and the Alliance for Open Media Patent License 1.0 are in THIRD-PARTY-NOTICES.TXT.
  • Registration: format HEIF, MIME types image/heif and image/avif, extensions .heif, .hif and .avif. A .heif or .hif file decodes only when it contains AV1. Files with HEVC content, for example most camera .hif files, throw ImageFormatException.

Decoder: still images, grids, animations and layered images. 8, 10 and 12 bits, monochrome to 4:4:4, lossless. Alpha, ICC, Exif, XMP, and the rotation, mirror, clean aperture and HDR properties.

Encoder (HeifEncoder): still images, animations and layered images. 8, 10 and 12 bits, monochrome to 4:4:4, lossy and lossless, alpha. Speeds Level0 to Level9, quality, tuning, rate control, film grain and tiling options. Writes ICC, Exif and XMP.

Benchmarks

Quality 75, speed 6, 4:2:0. ImageSharp runs on one thread. Magick (libheif with libaom) uses one thread per logical core. The SingleCore rows pin Magick to one core.

BenchmarkDotNet v0.15.8, Windows 11 (10.0.26300.9550)
AMD RYZEN AI MAX+ 395 w/ Radeon 8060S 3.00GHz, 1 CPU, 32 logical and 16 physical cores
.NET 10.0.12, X64 RyuJIT x86-64-v4, IterationCount=15, WarmupCount=5

Decode

Method TestImage Mean StdDev Ratio Allocated
'Magick Avif' Heif/Irvine_CA.avif 21.07 ms 0.623 ms 1.00 5.07 KB
'Magick Avif SingleCore' Heif/Irvine_CA.avif 20.21 ms 0.784 ms 0.96 5.07 KB
'ImageSharp Avif' Heif/Irvine_CA.avif 12.34 ms 0.115 ms 0.59 536.99 KB
'Magick Avif' Heif/libavif-kodim23-8b.avif 23.05 ms 0.891 ms 1.00 5.07 KB
'Magick Avif SingleCore' Heif/libavif-kodim23-8b.avif 21.75 ms 0.656 ms 0.94 5.07 KB
'ImageSharp Avif' Heif/libavif-kodim23-8b.avif 11.13 ms 0.210 ms 0.48 533.2 KB

Encode

Method TestImage Mean StdDev Ratio Allocated
'Magick Avif' Png/Bike.png 243.3 ms 3.64 ms 1.00 61.6 KB
'Magick Avif SingleCore' Png/Bike.png 440.1 ms 20.69 ms 1.81 70.07 KB
'ImageSharp Avif' Png/Bike.png 222.9 ms 1.21 ms 0.92 362.83 KB
'Magick Avif' Png/splash.png 157.6 ms 1.62 ms 1.00 63.51 KB
'Magick Avif SingleCore' Png/splash.png 268.8 ms 18.58 ms 1.71 62.01 KB
'ImageSharp Avif' Png/splash.png 156.8 ms 6.78 ms 1.00 648.72 KB

On one thread, ImageSharp decodes 1.7 to 2.1 times faster than Magick, and encodes in the same time as Magick on all 32 logical cores.

Native memory. The Allocated column counts managed memory only. An elevated run with the native memory profiler (07.10.2026, same machine) measured the native allocations per operation:

Operation Magick ImageSharp
Decode Irvine_CA.avif 13,408 KB 0 KB
Decode libavif-kodim23-8b.avif 12,420 KB 1,639 KB
Encode Bike.png 32,302 KB 1 KB
Encode splash.png 64,770 KB 23,408 KB

ImageSharp's native memory is its own unmanaged buffer pool, which keeps returned buffers for reuse.

@JimBobSquarePants

Copy link
Copy Markdown
Member

Oh wow! Thanks! 🙏 I’ll have a good dig through asap!!

@dlemstra

dlemstra commented Jan 4, 2024

Copy link
Copy Markdown
Member

I noticed that you used the x265 encoder. This is licensed under the GPL and cannot be used in this project.

@ynse01

ynse01 commented Jan 5, 2024

Copy link
Copy Markdown
Contributor Author

I noticed that you used the x265 encoder. This is licensed under the GPL and cannot be used in this project.

True, and that is why I used the Nokia repository as reference for my implementation, which has a license which I guess is compatible with ImageSharp. I did not start on writing an HVC Codec, due to patent concerns.

Nokia also includes a Patent waiver in its license. AVIF is supposed to be patent free, even its Codec. For AVIF I used SVT-AV1 as a reference, which has a 3-clause BSD license.

With all this, I hope the reference implementations that are mentioned now, can be used to create an AVIF Codec implementation in ImageSharp.

@JimBobSquarePants JimBobSquarePants left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I need some guidance with the licensing implications here.

@dlemstra ImageMagick uses libheif but has a custom license. How does that work?

/// <summary>
/// File Type.
/// </summary>
ftyp = 0x66747970U,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These will all need to use the correct casing.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK, went for automaticaly generated PascalCasing for now, if this needs to be nicer, can be done with third element.

Comment thread src/ImageSharp/Formats/Heif/Heif4CharCode.cs Outdated
Comment thread src/ImageSharp/Formats/Heif/HeifEncoderCore.cs Outdated
Comment thread src/ImageSharp/Formats/Heif/HeifItemLink.cs Outdated
@ynse01

ynse01 commented Mar 14, 2024

Copy link
Copy Markdown
Contributor Author

Not sure why one of the RowIntervalTests is failing, didn't touched that part of the code...

Comment thread src/ImageSharp/Formats/Heif/HeifDecoderCore.cs Outdated
Comment thread src/ImageSharp/Formats/Heif/HeifDecoderCore.cs Outdated
@ynse01

ynse01 commented Apr 1, 2024

Copy link
Copy Markdown
Contributor Author

@JimBobSquarePants Let me know if you still want to go ahead with this PR.

@JimBobSquarePants

Copy link
Copy Markdown
Member

@JimBobSquarePants Let me know if you still want to go ahead with this PR.

Hi @ynse01 I absolutely do want to continue but we'll need to get AV1 format together before I can merge the code. Thanks!

@ynse01

ynse01 commented Apr 2, 2024

Copy link
Copy Markdown
Contributor Author

@JimBobSquarePants Let me know if you still want to go ahead with this PR.

Hi @ynse01 I absolutely do want to continue but we'll need to get AV1 format together before I can merge the code. Thanks!

AV1 is quite complex with all its options, especially for implementing an efficient Encoder.

Would it be easier for you, if we focus first on getting the pure HEIF container support in, including all metadata (ICC, Exif and XMP) working ?
AV1 compression could than be a separate next step.

@JimBobSquarePants

Copy link
Copy Markdown
Member

My concern here would be merging an incomplete format. We’ll have to make all the parts internal if you want to do things in that order.

@ynse01

ynse01 commented Apr 4, 2024

Copy link
Copy Markdown
Contributor Author

I think I understand what you're trying to tell me:
From a user perspective, parsing only metadata out of an image is not considered supporting a format. Users expect pixels out of an image.
Which would mean I guess, that the minimal viable PR to be merged in would be something like a first image round-trip, using 1 chosen set of parameters. That would allow for adding the other (many) variations relatively quickly and keep up with user's expectations.

The downside I see with this approach, is that the internal design of the implementation doesn't get any feedback until the very end, making design changes a bit more painful. Not to mention the size of the first PR ...

Now that I know you would like to have the AVIF format supported, I'm motivated again to continue to implement it.

@JimBobSquarePants Let me know your thoughts on the approach, order and intermediate goals of AVIF support.

@JimBobSquarePants

Copy link
Copy Markdown
Member

I think I understand what you're trying to tell me: From a user perspective, parsing only metadata out of an image is not considered supporting a format. Users expect pixels out of an image. Which would mean I guess, that the minimal viable PR to be merged in would be something like a first image round-trip, using 1 chosen set of parameters. That would allow for adding the other (many) variations relatively quickly and keep up with user's expectations.

The downside I see with this approach, is that the internal design of the implementation doesn't get any feedback until the very end, making design changes a bit more painful. Not to mention the size of the first PR ...

Now that I know you would like to have the AVIF format supported, I'm motivated again to continue to implement it.

@JimBobSquarePants Let me know your thoughts on the approach, order and intermediate goals of AVIF support.

Yes exactly. If we ship with support for a format, it must support the full decode/encode process for one compression format. Our users expect that now as that's how we've always worked.

I would implement a single compression (the most common if you know that) algorithm then iterate once we know that is complete. Addition compression algorithms can be separate PRs.

Regarding reviewing the code. This is why it's good to have the PR open early so we can talk through changes as they're written. I try to read changes as soon as they appear in PRs but if there's ever anything you explicitly want me to look at just tag me.

I must say though, the quality of code you are delivering is very high!!

@ynse01

ynse01 commented Apr 6, 2024

Copy link
Copy Markdown
Contributor Author

Sounds like a plan !

The user community of AVIF encoders seems to be quite scattered, can't identify a dominant force there. Feel free to provide real-life images, so I can direct the support towards these.

For now, I plan for the following initial limitations to AV1 compression:

  • 8 bit depth only
  • Color pixels only
  • Base Profile for now
  • Square block sizes
  • Split partition type only, to keep the blocks square
  • DCT transform type only
  • Limited selection of Prediction modes
  • Limited subsampling options
  • No filters
  • No alpha
  • No sequences / animation

I do plan to take both lossy and lossless into account from the start.

@JimBobSquarePants In the meantime, you could help my local workflow, by adding the relevant file extensions to GIT LFS in Shared-infrastructure please:

*.heic              filter=lfs diff=lfs merge=lfs -text
*.hif               filter=lfs diff=lfs merge=lfs -text
*.avif              filter=lfs diff=lfs merge=lfs -text

@JimBobSquarePants

Copy link
Copy Markdown
Member

Sounds like a plan !

The user community of AVIF encoders seems to be quite scattered, can't identify a dominant force there. Feel free to provide real-life images, so I can direct the support towards these.

For now, I plan for the following initial limitations to AV1 compression:

  • 8 bit depth only
  • Color pixels only
  • Base Profile for now
  • Square block sizes
  • Split partition type only, to keep the blocks square
  • DCT transform type only
  • Limited selection of Prediction modes
  • Limited subsampling options
  • No filters
  • No alpha
  • No sequences / animation

I do plan to take both lossy and lossless into account from the start.

@JimBobSquarePants In the meantime, you could help my local workflow, by adding the relevant file extensions to GIT LFS in Shared-infrastructure please:

*.heic              filter=lfs diff=lfs merge=lfs -text
*.hif               filter=lfs diff=lfs merge=lfs -text
*.avif              filter=lfs diff=lfs merge=lfs -text

That all sounds great. Sorry about the delay, I've been very swamped.

If you use the following command in your fork you will get those changes.

git submodule foreach git pull origin main

Comment thread src/ImageSharp/Formats/Heif/HeifDecoderCore.cs Outdated
@brianpopow

Copy link
Copy Markdown
Collaborator

@ynse01 The Av1BitStreamReader seems to work as it should, I was able to parse Av1CodecConfiguration with it d8e7078. Maybe I can help with this PR, but I still in the process of reading the spec, there is a lot to grasp.

@ynse01

ynse01 commented Jun 24, 2024 •

Copy link
Copy Markdown
Contributor Author

@ynse01 The Av1BitStreamReader seems to work as it should, I was able to parse Av1CodecConfiguration with it d8e7078. Maybe I can help with this PR, but I still in the process of reading the spec, there is a lot to grasp.

There sure is a lot of grasp for AVIF.

The parsing of ObuSequenceHeader seems to work OK indeed, but somewhere in either ObuFrameHeader or in the handling of TileGroups there is a bug. The way AVIF (or technically Av1) intertwines BitStream coding with Symbol coding is pretty hard to debug. Merged in my latest changes even though some tests are still failing, as I did some minor refactoring.

@ynse01

ynse01 commented Jun 25, 2024

Copy link
Copy Markdown
Contributor Author

I fixed the bug :-) Tests are passing again.

@brianpopow There are a few TODO items in the Heif container area, if you like you could have a look at:

  • Item locations can be relative to multiple origins, this is incompletely implemented in the HeifLocation class.
  • Sorting of item locations, such that we can read the stream sequentially.
  • There is no connection between HeifDecoderCore and Av1Decoder classes. Ideally, this could be designed such that LegacyJpeg (and potentially Jpeg-XR) can be implemented similarly, thinking of an interface here. Look in 'ParseMediaData', which is currently hardcoded into LegacyJPEG.
  • We need code to split the incoming pixels into their respective Y, U and V channels, as Av1 uses the YUV pixel format exclusively.

I'll dive deeper into the Symbol coding and parsing the coefficients of the TileInfoBlocks.

@brianpopow

Copy link
Copy Markdown
Collaborator

I fixed the bug :-) Tests are passing again.

@brianpopow There are a few TODO items in the Heif container area, if you like you could have a look at:
...
I'll dive deeper into the Symbol coding and parsing the coefficients of the TileInfoBlocks.

Good job in finding the bug! I had some time on the weekend too look into your code. It's impressive howmuch is already done with the OBU reader. I thought maybe I could start with something small, maybe implementing parsing 5.9.14. Segmentation params as a start?

@ynse01

ynse01 commented Jun 25, 2024 •

Copy link
Copy Markdown
Contributor Author

I fixed the bug :-) Tests are passing again.
@brianpopow There are a few TODO items in the Heif container area, if you like you could have a look at:
...
I'll dive deeper into the Symbol coding and parsing the coefficients of the TileInfoBlocks.

Good job in finding the bug! I had some time on the weekend too look into your code. It's impressive howmuch is already done with the OBU reader. I thought maybe I could start with something small, maybe implementing parsing 5.9.14. Segmentation params as a start?

Great !
I see the implement-another-spec-section-virus has gotten to you also. If you like more of these, feel free to do so.
As a goal, we could support images that do not use the reduced still picture header feature and explicitly contain every OBU module out there. The test image Irvine_CA.avif is such an example, in ObuFrameHeaderTest you can uncomment it and work your way through the the OBU headers.

@ynse01

ynse01 commented Jul 4, 2024

Copy link
Copy Markdown
Contributor Author

@brianpopow I'm following the code path of the SVT-AV1 library, which seems to deviate quite a bit in the latest Syntax's. From Residual and down, I don't think the spec and the code map 1-to-1.
As the SVT-AV1 library has extensive intrinsic parts, I would like to leverage those and stick close to their implementation.

@brianpopow

Copy link
Copy Markdown
Collaborator

@brianpopow I'm following the code path of the SVT-AV1 library, which seems to deviate quite a bit in the latest Syntax's. From Residual and down, I don't think the spec and the code map 1-to-1. As the SVT-AV1 library has extensive intrinsic parts, I would like to leverage those and stick close to their implementation.

OK, thanks for the update, I was following the spec very closely. I will take a look into SVT-AV1 library. I have seen that I have implemented GetPlaneResidualSize() in 13173a3, which you already have done in a another place, that was a waste of time, I have overlooked that.

@brianpopow

Copy link
Copy Markdown
Collaborator

@brianpopow I'm following the code path of the SVT-AV1 library, which seems to deviate quite a bit in the latest Syntax's.

I tried to look into the SVT-AV1 repository, but I have some understanding issues, @ynse01 maybe you can help me understand what you are using as the reference code base. Just for clarification, when you say SVT-AV1 library you mean this repo: https://gitlab.com/AOMediaCodec/SVT-AV1 right? If so, I only see an encoder here, where is the decoder part?

@ynse01

ynse01 commented Jul 11, 2024 •

Copy link
Copy Markdown
Contributor Author

@brianpopow I'm following the code path of the SVT-AV1 library, which seems to deviate quite a bit in the latest Syntax's.

I tried to look into the SVT-AV1 repository, but I have some understanding issues, @ynse01 maybe you can help me understand what you are using as the reference code base. Just for clarification, when you say SVT-AV1 library you mean this repo: https://gitlab.com/AOMediaCodec/SVT-AV1 right? If so, I only see an encoder here, where is the decoder part?

Indeed, that is the repository that I've used as reference. The decoder has been removed recently, in this commit. As I've used an earlier version of that exact decoder as the basis for my code here. This is a bit of a worry from a maintenance point of view.

The SVT-AV1 library is still one of the more efficient encoders out there and its license is permissive. Hopefully we can learn a couple of tricks from them on the encoder.

@brianpopow

Copy link
Copy Markdown
Collaborator

Hi @ynse01, I still have some problems understanding howto decode a single image with the SVT-AV1 library.

I am using now the tagged version v2.1.0. I can see now there is an SvtAv1DecApp, but that seems to be for video files, if I am not mistaken. Is it not possible to decode just a single avif file? This would make comparing our implementation with the reference much easier, but I struggle to understand how this can be done.

@ynse01

ynse01 commented Jul 12, 2024

Copy link
Copy Markdown
Contributor Author

@brianpopow Running the decoder app on an AVIF file will not work. I don't think SVT-AV1 has the HEIF container parsing capabilities. It seems to use the IVF file format, but it supports more, see here. It appears to me that IVF can contain an AV1 bitstream although its meant for VP8. And yes, I know this whole container / bitstream and their naming is pretty confusing.

To summarize the terminology: Heif is the ISO base container (also known as ISOBMFF), where AV1 is just one of the supported bitstream formats. AVIF is to be interpreted here as an AV1 bitstream inside a Heif container, where the bitstream contains exactly one intra frame. With the same terminology, Section 1 of AVIF spec also touched on this briefly.

The workflow I used to "debug" SVT-AV1 is considerably less user friendly from what you suggest. I mainly used modified unit tests from their own test suite, to get the information that I need. Alternatively, when you convert the AVIF file into IVF, you could get it visualized in a debugger.

@brianpopow

brianpopow commented Jul 20, 2024 •

Copy link
Copy Markdown
Collaborator

@ynse01 I have noticed an issue in the bitstream reader in ReadLiteral(). When you read to the end of the stream the byte position would be advanced, resulting in an Index out of range issue. I have added a test demonstrating the issue: 62728b6

Also when there is only 4 bytes of data available, the initialization would also result in an index out of range issue, because the AdvanceToNextWord would be called twice in the constructor. I did noticed this when decoding example image jpeg444_xnconvert.avif when av1CodecConfiguration was parsed in HeifDecoderCore -> ParsePropertyContainer()

I have tried to adjust the ReadLiteral() implementation to fix that, but I did not want to override your code without asking for your opinion first. I have added a new Av1BitStreamReader2 to demonstrate this. It's adapted from libgav1, because I found the implementation of Svt-AV1 too complicated.

All Bitstream tests pass with that implementation.

@ynse01

ynse01 commented Jul 20, 2024

Copy link
Copy Markdown
Contributor Author

@brianpopow I like your new Av1BitStreamReader, for 3 reasons:

  • Fixes existing bug with reading large integers
  • Doesn't need extra padding bytes at the end
  • Has less jumps, which is more efficient with modern (Intel) CPU's

So, replaced the implementation. I did rework your suggested code a bit and implemented switching to Symbol reading again.

The last branch of the extreme overshoot test needs a running size below the
running target, which the first branch always takes. The intra error and
rate fields of the TPL block statistics were never written or read; only
external rate control reads them.
Parameters, locals and a few private fields named scratch now say what they
hold, for example intermediateRows, transposeStorage or readBuffer. The
chroma-from-luma guard message no longer names a reference function. Long
signatures and calls that the new names made longer are wrapped.
The high-bit-depth frame sets no previous source and updates no noise
estimate, because the superblock source comparison and the noise estimate
run only for 8-bit samples. Its resize path already matches the scaler that
frames above eight bits use.
Rename members, locals and the TPL partial file that used the word scratch.
Reword the comments so they say what each buffer holds.
Delete members with no references in source or tests, and the members
that only those members used.
Each inverse DCT and ADST operator now has one generic stage network over
a lane type. A static-abstract lane operator supplies the clamped add and
subtract, negation, butterfly and rounded multiply for the scalar,
128-bit and 256-bit widths. The scalar overload runs the same network.

The 32-point and 64-point networks run as stage-chain methods, so each
method keeps its lane arithmetic inside the JIT inlining budget. The
sparse operators call the shared later stages of the full operator.
Every caller passed true after the scalar overloads were removed.
The coding-pass search, the first pass and the OBMC search now run the same
full-sample diamond and mesh searches. Each caller supplies its error measure
as a struct cost type. The OBMC search keeps every stage of a repeated radius.
Replace the scalar DC, directional and intra-block-copy predictors with
independent references in the tests, move the byte pyramid fill operator
to its test, and delete the test-only DC block path, the lossy block
encoder entry and the scalar compound average with their tests.
…nsform

The vector 8x8 inverse doubled the identity row in sixteen-bit lanes with saturation and then halved it. An 8-bit
stream can code coefficients up to 32767, so values above 16383 came out as 16383. The doubling and the row shift
cancel, so the kernel now skips both for the identity row. A new test compares every inverse transform type and size,
at 8, 10 and 12 bits and at every vector tier, with a reference written from the AV1 specification.
The scale reads no samples, so it needs no sample operator and no operator member.
Nothing sets it after construction.
The diagnostic log path used a leading dot and backslashes, so it was never uploaded, and on Linux and macOS it became
a single file name. The script now builds the path with Join-Path inside tests/Images/ActualOutput. The failure
artifact also holds the blame sequence files, and its name includes the runner architecture, so the x64 and Arm Linux
jobs no longer write the same artifact.
@JimBobSquarePants

Copy link
Copy Markdown
Member

Well….. This works.

Each test writes the encoded animation to the test output folder, so the result can be played back.
… tests

The new input is leo.gif encoded with ffmpeg and libaom. It loops and keeps the source frame durations.
The GIF conversion tolerance now sits just above the measured difference.
The test results folder was never shown to hold the blame files on CI.
When the test process crashes, --blame writes a sequence file to the default dotnet test results folder. The file lists
the tests that ran before the crash. Comments now explain both uploaded folders.
Every caller now passes each argument. The values match the old defaults, so the encoder and decoder output does not change.
ResolveAv1Encoding takes its cancellation token last.
When a luma palette wins its own search but loses the full comparison, the retained intra mode now codes chroma with DC and uses
the transform size of the palette, as the search leaves them. The lookahead motion settings and the early speed settings now take
the configured tune and sharpness.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.