Skip to content

[DON'T MERGE] 0.15 → 1.0: user-facing changes not covered by the migration guide - #2121

Draft
AndreiKingsley wants to merge 1 commit into
masterfrom
report/user-facing-changes-0.15-to-1.0
Draft

AndreiKingsley wants to merge 1 commit into
masterfrom
report/user-facing-changes-0.15-to-1.0

Conversation

@AndreiKingsley

Copy link
Copy Markdown
Collaborator

Warning

DON'T MERGE. This PR has no file changes (one empty commit). It is a vehicle for a review report for #1630 (Improve 1.0 Migration guide).

What this is

This is an audit of every user-facing change between v0.15.0 (2024-12-09) and current master that is not yet fully described in MigrationTo_1_0.md. Use it as a checklist for the missing guide sections (Modules, Compiler plugin, JDBC, …) and for fixing the existing ones.

Scope and method

Sources reviewed:

  • 539 merged PRs (Add support for reading parquet file thanks to arrow-dataset #576 #577…[AIR] Fixed typos, grammar and misleading docs in KDocs and documentation #2116, merged 2024-12-09 → 2026-10-01). Each PR's title, body, linked issues, changed src/main files, and api/*.api dump changes were checked; diffs were read where impact was unclear.
  • 340 issues closed as completed in the same period. This includes compiler-plugin issues that were fixed in the Kotlin repo.
  • Release notes: 1.0.0-Beta2, Beta3, Beta4, Beta5, and 1.0.0-rc01.
  • Binary API dumps: a per-class diff of api/*.api for every module between v0.15.0 and master (core.api: +2876/−7030 lines).
  • Deprecations: every @Deprecated in */src/main/kotlin at v0.15.0 (42) vs master (1142: 719 WARNING, 147 ERROR, 276 HIDDEN).

Included: any public API change (new, changed, renamed, removed, or deprecated), and significant behavior or default changes. This also covers platform and dependency changes that users can see.

Excluded:

  • build, CI, tests, docs and examples;
  • pure bug fixes;
  • impl and @PublishedApi internal leaks;
  • KDoc-only declarations;
  • compiler-plugin annotations that do not change the API (@Interpretable, @Refine, …);
  • changes that are already fully covered by the migration guide.

All changes are relative to 0.15. Things that were added and removed again between betas are left out.

Legend

Mark Meaning
🆕 New API
🔄 Behavior/default change (same code, different result)
💥 Breaking change (source or binary incompatible, no deprecation cycle)
🗑️ Deprecation
❌ Removal

The Guide column is either — (not mentioned) or partial (with what is missing).

Summary

# Cluster Items Most important
0 Migration guide errata 10 Some recommended replacements are deprecated themselves or don't exist
1 Modules, artifacts & platform 13 dataframe-json / dataframe-jupyter split, Kotlin 2.4, Gradle/KSP plugin gone
2 Compiler plugin, schemas & codegen 14 codegen no longer camel-cases names, nested marker names, requireColumn
3 Access API deprecation 3 ~630 KProperty/accessor overloads are WARNING
4 DataFrame / column creation 15 toDataFrame() semantics, emptyDataFrame, ValueColumn<DataFrame>
5 Columns Selection DSL 7 exceptNew → except, Transformable* gone
6 Modification & reshaping 19 renameToCamelCase, merge, move/insert additions, new exceptions
7 GroupBy / Pivot 7 concatWithKeys, countDistinct, pivotCounts fix
8 Statistics 15 skipNA→skipNaN, median→Double, sum types, percentile
9 Types, parsing & conversion 8 UUID strings now parsed to Uuid by default
10 Schema, cast & comparison 4 schema == is order-sensitive, ComparisonMode
11 IO: general 4 format-guessing DataFrame.read() is ERROR; unified name repair
12 IO: CSV/TSV 5 parameter order, charset, Compression
13 IO: JSON 5 unifyNumbers, null arrays
14 IO: Excel 3 parseEmptyAsNull = true, no more duplicate-header exception
15 IO: Arrow / Parquet 5 readParquet, struct ↔ ColumnGroup, Instant written as Timestamp
16 IO: JDBC / SQL 14 the whole JDBC API; the guide section is a TODO
17 IO: Geo & OpenAPI 6 dataframe-geo-jupyter, OpenAPI is experimental
18 Rendering 11 print()/toString() output changed, format rework
19 Kotlin Notebook 5 GroupBy accessors, auto opt-ins
20 Utilities & misc 4 DataFrameError, none {}

0. Errata in the current migration guide

These are things the guide currently says that are wrong or misleading (checked against master sources).

Guide says Reality PRs
CSV table: CSV/TSV → CsvDeephaven/TsvDeephaven CsvDeephaven/TsvDeephaven are ERROR-deprecated themselves, together with the whole SupportedDataFrameFormat. Point users to readCsv/readTsv/writeCsv. #1916
"All outdated CSV IO functions raise WARNING in 1.0" DataFrame.read(delimiter = ..) and the CSV/TSV classes are ERROR. The old readDelimStr is HIDDEN and now resolves to the Deephaven one. #1033, #1057, #1916
df.toCSV(..), df.writeTSV(..) The real 0.15 names were toCsv(..) and writeCSV(..). There was no writeTSV. #1057
DataColumn.createFrameColumn(name, df, startIndices) → df.chunked(name, startIndices) The signature is df.chunked(startIndices, name = "groups"). The ReplaceWith also points to internal chunkedImpl. #1147
df.select { single { predicate } } → cols().filter { predicate }.single() cols() without arguments is ERROR-deprecated (#1369). Use all().filter { predicate }.single(). The ReplaceWith (SINGLE_PLAIN_REPLACE) has the same problem. #1176, #1369
cols() → all()/allCols() is presented as a regular deprecation It is ERROR already in 1.0 (no WARNING phase). #1369
df.convert { }.convertToDeprecatedInstant() / .convertToStdlibInstant() On Convert they are toDeprecatedInstant() / toStdlibInstant(). convertTo* exists only on DataColumn. #1368
df.convert { }.toURL() is ERROR "to be removed in 1.1" It is already removed. Only DataColumn.convertToURL() remains as ERROR. #1839
row.rowPercentile()/rowPercentileOrNull() listed as 0.15 API percentile did not exist in 0.15. It was added (#1060) and renamed (#1149) during the 1.0 cycle, so this row should be dropped. #1060, #1149
Big-numbers section: "median/percentile … some overloads throw" After #1482, all-columns df.median()/df.percentile() include big-number columns as plain comparables (no interpolation). df.cumSum() silently skips them (#1467). #1467, #1482

1. Modules, artifacts & platform

The guide section is a TODO.

Change Type PRs / issues Guide
JSON IO moved to the new dataframe-json module: readJson*, writeJson, toJson, toJsonWithMetadata, JSON, CustomEncoder, Base64ImageEncodingOptions. The package is unchanged. core no longer depends on kotlinx-serialization. The dataframe umbrella re-exports it, and dataframe-csv/dataframe-excel depend on it via api. Projects that depend only on dataframe-core must add dataframe-json. 💥 #1147, #100, #893 —
Jupyter integration moved to the new dataframe-jupyter module. %use dataframe loads it. dataframe-core ships libraries.json that auto-loads it when only core is on a notebook classpath. The minimum Kotlin Jupyter kernel was raised (release notes: use 0.16.0-736+). 💥 #1095, #1163, #1415, #775, #1140 —
dataframe-openapi was removed from the dataframe umbrella artifact (it was api in 0.15) and is now experimental. Add the dependency explicitly. 💥 #1115 —
Kotlin 2.0.20 → 2.4.20: kotlin-stdlib/kotlin-reflect 2.4 now come in transitively. Users need a compatible compiler, and the compiler plugin version always equals the Kotlin version. 💥 #1286, #1529, #1657, #1762, #1891, #2085 —
The KSP processor (symbol-processor-all) and the DataFrame Gradle plugin (org.jetbrains.kotlinx.dataframe, dataframes { schema {} }) are no longer published (last: 1.0.0-Beta4). The old alias id org.jetbrains.kotlin.plugin.dataframe now belongs only to the compiler plugin. ❌ #1400, #1670, #1843 — (only the separate "Migration from Gradle/KSP plugin" page)
JVM targets: core and IO modules still produce Java 8 bytecode. dataframe-jupyter, dataframe-geo-jupyter, and dataframe-openapi-generator require Java 11. 🔄 #1083, #1101, #700, #1100 —
Compiled with -jvm-default=no-compatibility: all *$DefaultImpls classes are gone. Binary-incompatible for Java callers and for code compiled against 0.15 that implements or calls DataFrame interfaces, so recompile. 💥 #1084, #1583, #1466 —
Many lambda-taking functions became inline (filter, count, all/any, first/last*, single*, map*, associate*, update.with, GroupBy.map, minBy/maxOf, …). Non-local return is now allowed, and the binary signatures changed. 💥 #1111, #1123, #1108, #1026 —
kotlinx-datetime 0.6.1 → 0.8.0-0.6.x-compat. In this version kotlinx.datetime.Instant/Clock are deprecated. 🔄 #1368, #1924 partial (the guide covers Instant but not the dependency bump)
dataframe-jdbc no longer brings the MariaDB driver transitively, so add org.mariadb.jdbc:mariadb-java-client yourself. The DuckDB driver is compileOnly too. ❌ #1267, #1366 —
dataframe-arrow now depends on arrow-dataset (native/JNI, needed for Parquet). Arrow went 18.1 → 19. 🔄 #577 —
dataframe-geo exposes GeoTools/JTS (gt-api, gt-main, gt-referencing, jts-core) as api, so you no longer need to redeclare them. 🔄 #2055, #1920 —
Android: Deephaven CSV ships META-INF/thirdparty-LICENSE. Android projects need packaging.resources.pickFirsts += "META-INF/thirdparty-LICENSE". 🔄 #1218, #1217 —

2. Compiler plugin, schemas & codegen

The guide section is a TODO.

Change Type PRs / issues Guide
Codegen packages renamed from org.jetbrains.dataframe.* to org.jetbrains.kotlinx.dataframe.*. This covers codeGen.{CodeGenerator, CodeGenResult, InterfaceGenerationMode} and keywords.{HardKeywords, SoftKeywords, ModifierKeywords}. 💥 #1095 —
generateInterfaces()/generateDataClasses() no longer convert names to camelCase: NameNormalizer.default is now the identity. NameNormalizer.from(..) was renamed to toCamelCaseFrom(..) with no deprecation. The old behavior is available through NameNormalizer.toCamelCaseByDelimiter. The same applies to dataframe-openapi-generator. 💥 #1922, #1921 —
Nested markers are named after their column (Members, License, …) instead of _DataFrameType1, and they are nested inside the parent. Configure this with the new nestedMarkerNameProvider: MarkerNameProvider parameter. 🔄 #1702 —
New and extended codegen API: df.generateDataClasses(..), DataFrameSchema.generateInterfaces/generateDataClasses(..), and new parameters extensionProperties, visibility, useFqNames, nameNormalizer. Old no-arg overloads are HIDDEN. 🆕 #1230, #1311, #1195 partial: the generateCode → generateInterfaces row omits that fields was removed and that extensionProperties now defaults to false (it was true in generateCode)
CodeWithConverter → CodeWithTypeCastGenerator (converter → typeCastGenerator: TypeCastGenerator, with() → declarationsWithCastExpression()). EMPTY_CONVERTER is deprecated in favor of TypeCastGenerator.Empty. 💥 #1263 —
The code generator no longer backtick-escapes soft keywords (val `value` → val value). 🔄 #1731, #1727 —
The experimental @Import annotation (@Import DataFrame.read(..)) was removed without deprecation. ❌ #1084 —
New df.requireColumn { "name"<Type>() }: a runtime check that also adds the column to the compile-time schema. Use it to migrate gradually from the String API. 🆕 #1715, #1808 —
New df.properties(): ColumnsScope<T>, which shows only generated column properties in completion. Generated accessors now use ColumnsScope as the receiver (binary change for precompiled accessors). 🆕 #957 —
@DisableInterpretation can now target functions, classes, and properties (for example inline fun <reified T> DataFrame<T>.foo() helpers). 🆕 #1790 —
String-API calls (select("a"), remove("a"), groupBy("a"), rename, statistics, …) are interpreted by the plugin (@StringApiInterpretable), so untyped frames get refined compile-time schemas. 🆕 #2057, #2071 —
Compiler plugin coverage grew a lot: split, gather, update, implode, fill*/drop*, convert (all), move/insert (incl. before, insideGroup), join, statistics on DataFrame/GroupBy, cumSum, distinct, parse, explode, allExcept, DataColumn.map*, valueCounts, toDataFrame/AddDsl/dataFrameOf builders, Java records, and more. 🆕 #1197, #1207, #1362, #1425, #1444, #1447, #1485, #1499, #1659, #1667, #1724, #1728, #1794, #1825, #1828, #1853, #1897, #2033 —
Wrong generated extension-property types now fail with a descriptive IllegalStateException (handleExtensionPropertyException) instead of a bare ClassCastException/NPE. 🔄 #1896, #1742 —
Schema extraction treats a DataFrame<T>? property as a value column, not a FrameColumn. 🔄 #1925, #1498 —

3. Access API (KProperty / column accessors) deprecation

The guide has no section for this yet.

Change Type PRs / issues Guide
About 630 @AccessApiOverload overloads that take KProperty, ColumnReference, or ColumnAccessor are deprecated at WARNING (DEPRECATED_ACCESS_API, kept across releases). This covers every operation, the CS DSL (col/colGroup/cols/all*/…), column(property), and row[colRef]. row[colRef] should become col[row], row["name"], or row.name. Migrate to the String API or extension properties. 🗑️ #1172, #1697, #1402, #1399, #2012, #1005, #1353 —
ColumnReference helpers withValues(..), gt, lt, eq, neq, length(), lowercase(), uppercase() are ERROR. 🗑️ #1412 —
ColumnReference<C>.map(infer) { } was made internal without an ERROR phase. columnGroup.map { } now resolves to DataFrame.map and returns a List. Use asDataColumn().map { } instead. 💥 #1396, #890 —

4. DataFrame / column creation

Change Type PRs / issues Guide
Iterable<T>.toDataFrame() creates a single value column for value-like types (primitives, String, enums, date-time, arrays, classes without properties). The typed overloads (Iterable<Int>.toDataFrame(): DataFrame<ValueProperty<Int>>, …) were removed, and the result is DataFrame<T>. 💥 #1081, #676 —
Map values are value types: toDataFrame()/unfold no longer expand them into size/entries/… columns. unfold uses the toDataFrame() rules. 🔄 #1097, #896 —
toDataFrame()/unfold now substitute generic type arguments (Pair<Int, String> → first: Int, second: String, not Any). excludeClasses works in nested levels, and empty nested classes no longer produce empty groups. 🔄 #1772, #1475, #1465, #1771 —
Java records are supported by toDataFrame(). 🆕 #1682, #1590 —
toDataFrame { } (CreateDataFrameDsl) no longer extends TraversePropertiesDsl. Top-level exclude()/preserve() are ERROR no-ops, so move them into properties { }. A new @CreateDataFrameDslMarker forbids calling outer-DSL functions inside properties { }. 💥 #2004, #1857 —
ValueColumn<DataFrame> is disallowed: creating one throws, and toDataFrame { preserve(DataFrame::class) } throws. Slicing a ValueColumn<AnyFrame?> without nulls yields a FrameColumn. ValueColumn.get(indices/range) now returns DataColumn<T>. 💥 #1856, #1928, #1852, #1926 —
emptyDataFrame<T>() is inline reified and returns a 0-row frame with T's columns (= DataFrame.emptyOf<T>()). Before, it had no columns. 💥 #1914, #463 —
DataColumn.empty(name) now returns DataColumn<Nothing> (type Nothing, previously Unit). New DataColumn.emptyOf<T>(name). 💥 #1551, #1546 —
Infix df concat otherDf (different type parameters) now returns DataFrame<Any>. Use the vararg/Iterable overloads or .cast<T>(). 💥 #1828 —
New dataFrameOf("a" to col1, "g" to columnOf("x" to col2)) and columnOf(vararg Pair<String, AnyBaseCol>), for hierarchical frames inline. 🆕 #1144, #353 —
New Iterable<Map<String, Any?>>.toDataFrame() (about 10× faster than manual) and Map<String, Any?>.toDataRow(). 🆕 #1463, #1635, #1433, #90, #719 —
New List<List<T>>.toDataFrame(header, containsColumns) in api. The io version is deprecated (WARNING). 🆕🗑️ #1486 —
New df.chunked(startIndices, name), toSequence()/toSequenceOf<T>(), and unfold(vararg roots, maxDepth). 🆕 #1147, #1046, #1127 partial (see errata for chunked)
DynamicDataFrameBuilder(checkDuplicateValues = true) no longer adds a renamed duplicate of an identical column. New get(name) and add(values, name). 🔄🆕 #1082, #715 —
dataFrameOf(..).withValues(..) with the wrong number of values throws a DataFrameError. 🔄 #1173 —

5. Columns Selection DSL

Change Type PRs / issues Guide
Experimental exceptNew became the stable except: colGroup.except { }, "g".except("a"). It is no longer infix, and @ExperimentalExceptCsDsl was removed. It keeps the group structure. The old accessor/ColumnsResolver overloads are ERROR and point to the { } form. 💥 #1030, #1038, #761, #932 —
CS DSL startsWith/endsWith (already ERROR in 0.15) were removed. Use nameStartsWith/nameEndsWith or colsName*. ❌ #1033, #993 —
TransformableColumnSet/TransformableSingleColumn (left over from recursively()) are internal. all, cols, colsOf, first, single, … now return plain ColumnSet/SingleColumn (binary change). 💥 #1302, #1037 —
ColumnsResolver.isSingleColumn()/isSingleColumnWithGroup() were removed without deprecation. ❌ #1076 —
New in-place rename inside the DSL: select { col(0) named "a" } / into "a". 🆕 #1666, #1199 —
The ColumnsContainer.mapToColumn(name, type: KType, …) overload became internal. A new DataFrame.mapToColumn(..) keeps df.mapToColumn working. ❌ #1360 partial (only mapToColumn → expr)
New predicate-less colsAtAnyDepth(), colsInGroups(), single(), singleCol(). 🆕 #1176, #1369 partial (see errata)

6. Modification & reshaping operations

Change Type PRs / issues Guide
renameToCamelCase() was rewritten. It splits on any non-alphanumeric character and on case and digit boundaries, then lowercases words: ORDER_DATE → orderDate (was oRDERDATE). DELIMITERS_REGEX/DELIMITED_STRING_REGEX were removed, and ColumnReference.renameToCamelCase() is new. 🔄 #1072, #1029, #988 —
merge { }.by(..)/asStrings() return the new MergeWithTransform, and notNull() narrows types. 💥 #1052 —
New move { }.before { }, insert(col).before { }, and insideGroup overloads (toStart(insideGroup), toEnd(..), to(index, ..), moveTo*(insideGroup)), plus ColumnsWithDifferentParentException. 🆕 #1474, #1517, #1489, #1021, #1516, #1255 partial (only toLeft/toRight renames)
insert(col).under { path } now creates missing parent groups. 🔄 #1739, #1411 partial (only the under(path) deprecation)
Insert failures throw the new sealed InsertException (NotAColumnGroupInsertException, DuplicateColumnPathInsertException). GroupBy.values() explains key/group column clashes. 🆕🔄 #1905, #1569 —
ungroup throws UngroupWrongColumnKindException for non-group columns. explode throws ExplodeWrongColumnKindException for columns that are not List or frame columns (before, both silently did nothing). explode(verify = false) restores the lenient behavior. 🔄 #1847, #295 —
split { }.by(..).into(names) always creates all named columns and fills missing values with null/default. 🔄 #1776, #1758 —
Exceptions inside update { }.with { } (and fill*) are wrapped in IllegalStateException("Could not update column '<name>'"). 🔄 #1551 —
append: null into a FrameColumn adds an empty frame, append() with no values returns the receiver, and a wrong value count throws IllegalArgumentException (was IllegalStateException). 🔄 #2061, #1927, #1952 —
Gather, Corr, and SplitWithTransform are no longer data classes (copy/componentN/equals are gone). 💥 #1182, #1826 —
ColumnMatch (a match b in join) changed from a class to an interface, with a ColumnMatch(left, right) factory function. 💥 #1139 —
The JoinType.allowLeftNulls/allowRightNulls extensions became internal without deprecation. ❌ #1574 —
New typed gather { listCols }.explodeLists(), so no cast is needed. 🆕 #1269 —
DataRow.explode(..): named parameter selector → columns. 💥 #1291 —
AddDsl ("col" from { }, expr { } into, …) receives AddExpression, so prev()?.newValue() works. 🆕 #1303, #614 —
DataRow<T>.transpose() → AnyRow.transpose() (the type parameter was dropped). 💥 #1553 —
df.filterNotNull { } was added as an ERROR pointer to dropNulls { } (for discoverability). 🗑️ #1334, #1099 —
Sort DSL: col.reverse() inside sortBy { } is now ERROR and points to reversed(). 🗑️ #1936, #1526 partial (only desc() → reversed())
unfold { } has new roots/maxDepth parameters (binary change). 💥 #1127 —

7. GroupBy / Pivot / aggregation

Change Type PRs / issues Guide
New groupBy.concatWithKeys(), which keeps the key columns. 🆕 #1107, #673 —
New groupBy.countDistinct() / countDistinct { cols }. 🆕 #1875, #533 —
pivotCounts inside aggregate { } now counts. Before, it returned pivotMatches booleans. 🔄 #1554, #1552 —
pivot/pivotCounts/pivotMatches shortcuts take a PivotDsl selector, so then works. 💥 #1572, #1548 —
df.asGroupBy { } moves the group column to the end, so aggregation results have a predictable column order. 🔄 #1110 —
PivotGroupBy.maxOf { }/minOf { } produce null for empty groups instead of throwing. 🔄 #2012 —
New GroupBy<*, *>.cast<T, G>(). 🆕 #1263 —

8. Statistics

The guide covers only BigDecimal/BigInteger and the row*Of renames.

Change Type PRs / issues Guide
skipNA → skipNaN in sum/mean/std (all For/Of/row variants on column, frame, GroupBy, Pivot), so named calls mean(skipNA = true) no longer compile. skipNaN was added to min/max/median/percentile. cumSum keeps skipNA. 💥 #1108, #1119, #1122, #1149 —
median of number columns returns an interpolated Double (quantile R8). In 0.15 it returned the column type with integer division. Other Comparable columns return the lower middle value (R3). New medianBy/medianByOrNull. 🔄 #1122, #1149, #566 partial
New percentile(p) family (percentile*, percentileFor/Of/By, rowPercentileOf), with the same typing as median. 🆕 #1060, #1149, #543 partial (see errata)
sum result types: Byte/Short → Int, mixed Number → unified type, empty input → typed 0. 🔄 #1103, #961 —
mean always returns Double (NaN for empty input), and DataColumn.meanOrNull() was removed. std uses ddof (default 1). 💥 #1091, #1119 —
Number unification: mixed-Number columns are converted to the smallest common type (Int + Float → Double), which fixes rowSum. sum()/mean()/std() without columns select only primitive or mixed-number columns. 🔄 #1070, #1078, #1068 —
min/max/minBy/maxOf/…: bound is T : Comparable<T & Any>?, many are inline reified, and mixed-number columns are no longer supported. 💥 #1108 partial (only the rowMin/rowMax rename)
minFor/maxFor/medianFor/percentileFor accept columns of different comparable types (Comparable<*>?). 🆕 #1430, #1429 —
rowXOf<T>() selects row values by runtime value type, not by column type. 🔄 #1119 —
describe() has new p25/p75 columns (ColumnDescription.p25/p75) around median. 🔄 #1060, #1054 —
cumSum follows the sum typing: Byte/Short → Int, and nullable Double/Float → non-null with NaN. 🔄 #1152 —
AnyCol.isBigNumber() was removed without deprecation. New isPrimitiveNumber(), isMixedNumber(), isPrimitiveOrMixedNumber(). ❌🆕 #1091, #1103 —
Public stats helpers for Kotlin collections in ...dataframe.math (Iterable.mean(), std, median, varianceAndMean, BasicStats) were made internal without deprecation. ❌ #1065 —
digitize was removed (all overloads) without deprecation. ❌ #1125, #1124 —
Big numbers in all-columns median/percentile/cumSum: see errata. 🔄 #1467, #1482 partial

9. Types, parsing & conversion

Change Type PRs / issues Guide
parseExperimentalUuid now defaults to true: parse(), convertTo, and readCsv/readTsv/… turn UUID-formatted strings into kotlin.uuid.Uuid. Opt out with ParserOptions(parseExperimentalUuid = false) or DataFrame.parser.parseExperimentalUuid = false. The option and UUID parsing itself are new since 0.15. 🔄🆕 #1287, #1306, #1891, #1006 —
useFastDoubleParser now defaults to true, with a fallback that also tries alternative decimal symbols for the locale (more lenient). 🔄 #1040, #1039 —
Char columns: df.parse() parses them ('1' → Int, 'T' → Boolean), convert { }.to<>()/convertTo<>() fall back to String (for example to enums), and DataColumn<Char?>.parse()/tryParse() are new. 🔄🆕 #1420, #998 —
New convert shortcuts: toJavaLocalDate/LocalTime/LocalDateTime/Instant/Duration (with pattern/DateTimeFormatter + locale), toDuration, toYearMonth, toUtcOffset, toDateTimeComponents(..), toLocalDate(format: DateTimeFormat), … convertTo<T>()/convert.to<T>() accept parserOptions. 🆕 #1777, #1839, #876 partial (only toLocalDate(pattern)/toJavaLocalDate)
Kotlin date-time convert overloads with a locale (9 functions) and DataFrame.parser.addDateTimePattern(..) are already ERROR. 🗑️ #1777 partial (replacements shown, level not stated)
ParserOptions.dateTimeFormatter/dateTimePattern properties were removed (not only the constructor parameters), and copy() changed. New exceptions: ColumnTypeMismatchesColumnValuesException and TypeConversionException.extraInformation. 💥 #1777 partial
convert shortcut overloads were merged into nullable-receiver variants, and the non-null Convert<T, Any>.toInt()/toStr()/… overloads were removed (source-compatible, binary-incompatible). 💥 #1118, #1839, #1695 —
String → URL conversion uses URI(str).toURL(), so malformed URLs (for example with spaces) now throw. 🔄 #1175 —

10. Schema, cast & comparison

Change Type PRs / issues Guide
DataFrameSchema.equals() takes column order into account (also for nested groups and frames). Use schema.compare(other).matches() for an order-insensitive comparison. 🔄 #1505, #2042, #1502 partial (only Equals → Matches)
compare(other, strictlyEqualNestedSchemas) → compare(other, comparisonMode = ComparisonMode.LENIENT) (LENIENT/STRICT/STRICT_FOR_NESTED_SCHEMAS). ColumnSchema is sealed, and CompareResult.plus is new. 💥 #1252, #1154, #1222 —
compileTimeSchema() is ordered like schema() and has a new ordered parameter. 🔄🆕 #990, #1427 —
changeType(type) is a public member of ValueColumn (it was on the internal DataColumnInternal). 🆕 #1750, #1675 —

11. IO: general

Change Type PRs / issues Guide
Format-guessing readers are ERROR: DataFrame.read(file/url/path/String), DataRow.read(..), File/URL/Path.readDataFrame()/readDataRow(), the whole SupportedDataFrameFormat service (CsvDeephaven, TsvDeephaven, Excel, ArrowFeather, Jdbc, JSON read overrides), buildCodeForDB, and CodeGenerator.urlDfReader. Use readCsv/readJson/readExcel/readParquet/readSqlTable/… 🗑️ #1916, #450, #705 partial (only read(delimiter = ..); see errata)
Column-name repair is unified: every reader de-duplicates with ColumnNameGenerator (a, a1, a2). readExcel no longer throws DuplicateColumnNamesException, and JDBC gives id1 instead of id_1. The deprecated readExcel(nameRepairStrategy = ..) overloads ignore the argument, so just drop it. 🔄 #1904, #1632, #387 partial (no naming scheme, no "Excel stops throwing", no "drop the argument")
java.nio.file.Path overloads for all IO (read, readJson/writeJson, readExcel/writeExcel, Arrow, Geo, CSV, importDataSchema). 🆕 #1563, #527 —
Custom SupportedDataFrameFormats must implement readDataFrame(path: Path, ..) (the File variant now has a default). isOpenApi(File) → isOpenApi(Path). 💥 #1563 —

12. IO: CSV/TSV

Change Type PRs / issues Guide
Parameter order changed in readCsv/readTsv/readDelim: compression moved to the end, and a new charset: Charset? comes after header (the BOM is auto-detected, UTF-8 fallback). Old overloads are HIDDEN, so positional calls break and named ones keep working. 💥🆕 #1142, #1665, #1141, #1555 —
The Deephaven reader's own date-time inference is disabled, so date-times are parsed by DataFrame's parser (offset strings become Instant instead of a shifted LocalDateTime). 🔄 #1057, #1047 —
Compression moved from dataframe-csv to core (same package) and became a fun interface (Gzip/Zip/None objects, Compression.of(..)). Compression.Custom(..) was removed, so use the SAM form. 💥 #1057 —
The old Apache readDelim(InputStream/Reader) and CSVType are WARNING. The old readDelimStr is HIDDEN, so calls resolve to the Deephaven one. 🗑️ #1057 partial (only the generic sentence)
InputStream.skippingBomCharacters() also strips UTF-16/32 BOMs. 🔄 #1665 —

13. IO: JSON

Change Type PRs / issues Guide
unifyNumbers = true on all readJson* and on JSON(..): mixed numbers get a common type instead of Number, and Float can be inferred. Inferred types may change. 🔄🆕 #1073, #557 —
readJson(..): null where an array is expected gives null (List<T>?) instead of an empty list. toJson writes null lists as null, and a user column named value/array no longer breaks writing. 🔄 #2037, #2056, #2035, #2045, #2046, #2048 —
readJson(keyValuePaths = ..) produces name/value columns (was key/value). 🔄 #1545 —
@JsonOptions(typeClashTactic = ..) takes the String constants JsonOptions.TypeClashTactics.ARRAY_AND_VALUE_COLUMNS/ANY_COLUMNS instead of the enum. 💥 #1147 —
New readJson(stream, .., format: Json) for a custom kotlinx-serialization Json (lenient parsing, comments, …). 🆕 #1869, #459 —

14. IO: Excel

Change Type PRs / issues Guide
New readExcel(.., parseEmptyAsNull = true): empty-string cells become null, which changes inferred types and nullability. 🔄 #1247, #1166 —
writeExcel to a new file uses the streaming SXSSFWorkbook (lower memory). keepFile = true with a missing or empty file creates a workbook instead of failing. 🔄 #1075, #1016 —
Duplicate headers are renamed instead of throwing (see IO: general). 🔄 #1904 partial

15. IO: Arrow / Parquet

Change Type PRs / issues Guide
New DataFrame.readParquet(urls/paths/files, nullability, batchSize) (Arrow Dataset, native). ARROW_PARQUET_DEFAULT_BATCH_SIZE is public. 🆕 #577, #576 —
StructVector ↔ ColumnGroup: structs are read as column groups (before, value columns of map-like values), and groups are written as structs. Nullable structs respect the null bit (children become nullable). 🔄 #1621, #2043, #536, #2041 —
ListVector/LargeListVector of primitives are read as List<T>, and lists of structs as FrameColumn. 🆕 #1807, #1804 —
Timestamps with a time zone are read as kotlin.time.Instant (before, NotImplementedError). Instant columns are written as Timestamp(MICROSECOND, "UTC") instead of Utf8 strings, and Timestamp targets are pinned to UTC. 🔄 #2063, #926 —
New ConvertingMismatch.PrecisionReduced/ValueOutOfRange subclasses, so an exhaustive when over the sealed class is source-incompatible. 💥 #2063 —

16. IO: JDBC / SQL databases

The guide section is a TODO, and almost none of this went through a deprecation cycle.

Change Type PRs / issues Guide
Schema API renamed without aliases: DataFrame.getSchemaForSqlTable/ForSqlQuery/ForResultSet/ForAllSqlTables(..) → DataFrameSchema.readSqlTable/readSqlQuery/readResultSet/readAllSqlTables(..), and Connection/DbConnectionConfig/ResultSet.getDataFrameSchema(..) → readDataFrameSchema(..). 💥 #1487 —
limit: Int = Int.MIN_VALUE → limit: Int? = null in all readers. A new trailing configureStatement: (PreparedStatement) -> Unit. The old signatures were removed. 💥 #1487, #680 —
New javax.sql.DataSource overloads for every reader and schema reader (connection pools). 🆕 #1487, #1424 —
Default value mapping: TIMESTAMP → kotlin.time.Instant (was java.sql.Timestamp), Java LocalDateTime → kotlinx LocalDateTime, UUID → kotlin.uuid.Uuid, MySQL/MariaDB BIGINT UNSIGNED → BigInteger (BigInteger statistics are not supported), MariaDB INT UNSIGNED → Long. PostgreSQL extension types are supported. 🔄 #1632, #1735, #1588, #461, #762 —
DbConnectionConfig(.., readOnly = true): connections from a config are read-only by default (auto-commit off, rolled back after reading). It is no longer a data class (componentN is gone), and toString() masks the password. 💥🔄 #1325, #1383, #1308, #1365 —
SQL validation: new validation: SqlValidation = SqlValidation.None on readSqlQuery/readDataFrame, so queries are not validated by default. In 0.15 they had to start with SELECT and contain no ;. SqlValidation.ReadOnly allows only one SELECT/WITH/VALUES/TABLE/EXPLAIN statement. Table names in readSqlTable are always validated. 🔄🆕 #1117, #1324, #1327, #1881, #1307, #1671 —
DbType overhaul (custom DbTypes must migrate): convertSqlTypeToColumnSchemaValue/convertSqlTypeToKType/extractValueFromResultSet were removed in favor of getExpectedJdbcType → getValueFromResultSet → preprocessValue → getTargetColumnSchema/buildDataColumn. sqlQueryLimit → buildSqlQueryWithLimit. New hooks: quoteIdentifier, configureReadStatement, createConnection, defaultFetchSize, defaultQueryTimeout, tableTypes, getTableColumnsMetadata. 💥 #1632, #1487, #1283, #1573, #1273, #1587 —
New AdvancedDbType and the JdbcToDataFrameConverter pipeline (DbResultSetReader, DbValuePreprocessor, DbColumnBuilder, jdbcToDfConverterFor, with*). 🆕 #1632, #462 —
TableColumnMetadata/TableMetadata moved from ...dataframe.io to ...dataframe.io.db. 💥 #1487 —
New built-in DuckDb (auto-detected from jdbc:duckdb:): STRUCT → column group, LIST/ARRAY, MAP, JSON, GEOMETRY, TIME_NS, VARIANT, and read-only mode. 🆕 #1366, #1632, #1924, #1341, #537 —
H2(dialect: DbType = MySql) → H2(mode: H2.Mode = H2.Mode.Regular). The default is no longer the MySQL dialect, and H2(dialect)/H2.MODE_* are deprecated. URLs without MODE= are supported. 💥 #1598, #1595 —
Sqlite changed from an object to a class (AdvancedDbType): use Sqlite.default or Sqlite(), plus Sqlite.withCustomConverters { forType<..>(..) {}; forColumn<..>(..) {} }. Declared types are mapped: BOOLEAN → Boolean, DATE/DATETIME/TIME/TIMESTAMP → kotlinx date-time/Instant. 💥🔄 #1741, #2001, #2002, #964, #1013, #1747 —
readAllSqlTables without catalogue falls back to connection.catalog (MySQL/MariaDB/MSSQL: only the URL's database). tableTypes defaults to TABLE, BASE TABLE. 🔄 #1901, #1283, #1746 —
Duplicate result-set column names: id, id1 (was id, id_1). 🔄 #1632, #1904 partial

17. IO: Geo & OpenAPI

Change Type PRs / issues Guide
New artifact dataframe-geo-jupyter. %use dataframe(enableExperimentalGeo=true) no longer loads Geo, so use %use dataframe-geo. dataframe-geo no longer depends on dataframe-jupyter. 💥 #1361, #1155 —
GeoDataFrame.modify { } also receives the frame as it. 🆕 #1086, #1049 —
readShapefile(dir) finds <dir>.shp inside a directory. writeGeoJson writes ordered feature IDs, so row order is preserved. 🆕🔄 #1641, #1640, #1567, #1565 —
OpenAPI support is experimental and opt-in: @ImportDataSchema(enableExperimentalOpenApi = true), or %use dataframe(enableExperimentalOpenApi=true) in notebooks. 🔄 #1115 —
AdditionalProperty<T> extends NameValueProperty: key → name (and the generated column is name). 💥 #1545 —
The OpenAPI generator keeps original property names (no camelCase). 🔄 #1922 —

18. Rendering: HTML, format, print, Markdown

Change Type PRs / issues Guide
print()/renderToString()/DataFrame.toString() output changed: column types are shown by default (columnTypes = true, on a separate row), an "N columns × M rows" footer appears on truncation, and DataColumn.toString() includes the type. String-snapshot tests will break. 🔄 #1760 —
print/renderToString: rowsLimit/valueLimit are Int? (null means unlimited), and there are new borderStyle: StringBorderStyle and rowIndex parameters. renderToString is public. The title changed from Data Frame [..] to DataFrame [..]. 🆕🔄 #1760, #1395, #1364 —
New renderToMarkdown(..) and MarkdownAlignment. 🆕 #1760, #525 —
format { } works on nested columns and column groups. RowColFormatter (and DisplayConfiguration.cellFormatter) receives ColumnWithPath<C> instead of DataColumn<C>. 💥🆕 #1374, #1898, #1356 —
Chained format { }.where { }.where { } now combine with AND (the last one used to win). textColor also applies to numbers and nulls. 🔄 #1346, #1280 —
Formatting no longer leaks into nested FrameColumn frames with same-named columns. Format nested frames explicitly. 🔄 #1443, #982 —
New format helpers: at(rows), notNull()/notNull { }, FormattedFrame.format(vararg String). 🆕 #1346 —
New formatHeader { cols }.with { } and DisplayConfiguration.headerFormatter. The FormatClause/FormattedFrame constructors changed. 🆕 #1459, #1470 —
FormattedFrame.toHTML/toStandaloneHTML were removed without deprecation, so use toHtml/toStandaloneHtml (now with cellRenderer and getFooter). 💥 #1175, #1892 partial (only DataFrame.toHTML)
FormattingDSL/RGBColor are now typealiases (the old classes are gone from bytecode, so it is a binary change). 💥 #1346 partial (the source rename is covered)
FormattedFrame renders formatted in Kotlin Notebook (toJsonWithMetadata(isFormatted)). 🆕 #1405, #1354 —

19. Kotlin Notebook / Jupyter

Change Type PRs / issues Guide
Typed accessors are generated for GroupBy variables (for example groupBy.keys.k in the next cell). 🆕 #1263, #1221 —
Generated accessors automatically opt in to ExperimentalTime/ExperimentalUuidApi for Instant/Uuid columns. 🔄 #1601, #1602, #1822 —
A ColumnGroup cast to DataFrame renders as a frame of its sub-columns. 🔄 #1460, #1245 —
Notebook markers have fields, so the compiler plugin can read schemas from earlier cells. 🔄 #1154, #867 —
Faster column sorting in the table view, including sorting List/DataFrame columns by size. 🔄 #1639, #1436 —

20. Utilities & misc

Change Type PRs / issues Guide
New DataFrame.none { } and DataColumn.none { }. 🆕 #1462, #1431 —
New marker interface exceptions.DataFrameError, implemented by ColumnNotFoundException, DuplicateColumnNamesException, and ColumnsWithDifferentParentException. 🆕 #1173, #1489 —
KeyValueProperty → NameValueProperty: the generated key column/property extensions are deprecated in favor of name. 🗑️ #1545, #659 partial (only the interface rename)
NamedValue.copy(..) is no longer public. ❌ #2014 —

Notable things deliberately left out

🤖 Generated with Claude Code

…the migration guide

Empty commit; the report is in the PR description (see #1630).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant