[DON'T MERGE] 0.15 → 1.0: user-facing changes not covered by the migration guide - #2121
Draft
AndreiKingsley wants to merge 1 commit into
Draft
AndreiKingsley wants to merge 1 commit into
AndreiKingsley wants to merge 1 commit into
Conversation
…the migration guide Empty commit; the report is in the PR description (see #1630). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Warning
DON'T MERGE. This PR has no file changes (one empty commit). It is a vehicle for a review report for #1630 (Improve 1.0 Migration guide).
What this is
This is an audit of every user-facing change between
v0.15.0(2024-12-09) and currentmasterthat is not yet fully described inMigrationTo_1_0.md. Use it as a checklist for the missing guide sections (Modules, Compiler plugin, JDBC, …) and for fixing the existing ones.Scope and method
Sources reviewed:
src/mainfiles, andapi/*.apidump changes were checked; diffs were read where impact was unclear.api/*.apifor every module betweenv0.15.0andmaster(core.api: +2876/−7030 lines).@Deprecatedin*/src/main/kotlinatv0.15.0(42) vsmaster(1142: 719WARNING, 147ERROR, 276HIDDEN).Included: any public API change (new, changed, renamed, removed, or deprecated), and significant behavior or default changes. This also covers platform and dependency changes that users can see.
Excluded:
impland@PublishedApi internalleaks;@Interpretable,@Refine, …);All changes are relative to 0.15. Things that were added and removed again between betas are left out.
Legend
The Guide column is either — (not mentioned) or partial (with what is missing).
Summary
dataframe-json/dataframe-jupytersplit, Kotlin 2.4, Gradle/KSP plugin gonerequireColumnKProperty/accessor overloads areWARNINGtoDataFrame()semantics,emptyDataFrame,ValueColumn<DataFrame>exceptNew→except,Transformable*gonerenameToCamelCase,merge,move/insertadditions, new exceptionsconcatWithKeys,countDistinct,pivotCountsfixskipNA→skipNaN,median→Double,sumtypes,percentileUuidby defaultschema ==is order-sensitive,ComparisonModeDataFrame.read()isERROR; unified name repaircharset,CompressionunifyNumbers,nullarraysparseEmptyAsNull = true, no more duplicate-header exceptionreadParquet, struct ↔ColumnGroup,Instantwritten as Timestampdataframe-geo-jupyter, OpenAPI is experimentalprint()/toString()output changed,formatreworkDataFrameError,none {}0. Errata in the current migration guide
These are things the guide currently says that are wrong or misleading (checked against
mastersources).CSV/TSV→CsvDeephaven/TsvDeephavenCsvDeephaven/TsvDeephavenareERROR-deprecated themselves, together with the wholeSupportedDataFrameFormat. Point users toreadCsv/readTsv/writeCsv.WARNINGin 1.0"DataFrame.read(delimiter = ..)and theCSV/TSVclasses areERROR. The oldreadDelimStrisHIDDENand now resolves to the Deephaven one.df.toCSV(..),df.writeTSV(..)toCsv(..)andwriteCSV(..). There was nowriteTSV.DataColumn.createFrameColumn(name, df, startIndices)→df.chunked(name, startIndices)df.chunked(startIndices, name = "groups"). TheReplaceWithalso points to internalchunkedImpl.df.select { single { predicate } }→cols().filter { predicate }.single()cols()without arguments isERROR-deprecated (#1369). Useall().filter { predicate }.single(). TheReplaceWith(SINGLE_PLAIN_REPLACE) has the same problem.cols()→all()/allCols()is presented as a regular deprecationERRORalready in 1.0 (noWARNINGphase).df.convert { }.convertToDeprecatedInstant()/.convertToStdlibInstant()Convertthey aretoDeprecatedInstant()/toStdlibInstant().convertTo*exists only onDataColumn.df.convert { }.toURL()isERROR"to be removed in 1.1"DataColumn.convertToURL()remains asERROR.row.rowPercentile()/rowPercentileOrNull()listed as 0.15 APIpercentiledid not exist in 0.15. It was added (#1060) and renamed (#1149) during the 1.0 cycle, so this row should be dropped.median/percentile… some overloads throw"df.median()/df.percentile()include big-number columns as plain comparables (no interpolation).df.cumSum()silently skips them (#1467).1. Modules, artifacts & platform
The guide section is a TODO.
dataframe-jsonmodule:readJson*,writeJson,toJson,toJsonWithMetadata,JSON,CustomEncoder,Base64ImageEncodingOptions. The package is unchanged.coreno longer depends on kotlinx-serialization. Thedataframeumbrella re-exports it, anddataframe-csv/dataframe-exceldepend on it viaapi. Projects that depend only ondataframe-coremust adddataframe-json.dataframe-jupytermodule.%use dataframeloads it.dataframe-coreshipslibraries.jsonthat auto-loads it when only core is on a notebook classpath. The minimum Kotlin Jupyter kernel was raised (release notes: use0.16.0-736+).dataframe-openapiwas removed from thedataframeumbrella artifact (it wasapiin 0.15) and is now experimental. Add the dependency explicitly.kotlin-stdlib/kotlin-reflect2.4 now come in transitively. Users need a compatible compiler, and the compiler plugin version always equals the Kotlin version.symbol-processor-all) and the DataFrame Gradle plugin (org.jetbrains.kotlinx.dataframe,dataframes { schema {} }) are no longer published (last: 1.0.0-Beta4). The old alias idorg.jetbrains.kotlin.plugin.dataframenow belongs only to the compiler plugin.coreand IO modules still produce Java 8 bytecode.dataframe-jupyter,dataframe-geo-jupyter, anddataframe-openapi-generatorrequire Java 11.-jvm-default=no-compatibility: all*$DefaultImplsclasses are gone. Binary-incompatible for Java callers and for code compiled against 0.15 that implements or calls DataFrame interfaces, so recompile.inline(filter,count,all/any,first/last*,single*,map*,associate*,update.with,GroupBy.map,minBy/maxOf, …). Non-localreturnis now allowed, and the binary signatures changed.kotlinx-datetime0.6.1 →0.8.0-0.6.x-compat. In this versionkotlinx.datetime.Instant/Clockare deprecated.Instantbut not the dependency bump)dataframe-jdbcno longer brings the MariaDB driver transitively, so addorg.mariadb.jdbc:mariadb-java-clientyourself. The DuckDB driver iscompileOnlytoo.dataframe-arrownow depends onarrow-dataset(native/JNI, needed for Parquet). Arrow went 18.1 → 19.dataframe-geoexposes GeoTools/JTS (gt-api,gt-main,gt-referencing,jts-core) asapi, so you no longer need to redeclare them.META-INF/thirdparty-LICENSE. Android projects needpackaging.resources.pickFirsts += "META-INF/thirdparty-LICENSE".2. Compiler plugin, schemas & codegen
The guide section is a TODO.
org.jetbrains.dataframe.*toorg.jetbrains.kotlinx.dataframe.*. This coverscodeGen.{CodeGenerator, CodeGenResult, InterfaceGenerationMode}andkeywords.{HardKeywords, SoftKeywords, ModifierKeywords}.generateInterfaces()/generateDataClasses()no longer convert names to camelCase:NameNormalizer.defaultis now the identity.NameNormalizer.from(..)was renamed totoCamelCaseFrom(..)with no deprecation. The old behavior is available throughNameNormalizer.toCamelCaseByDelimiter. The same applies todataframe-openapi-generator.Members,License, …) instead of_DataFrameType1, and they are nested inside the parent. Configure this with the newnestedMarkerNameProvider: MarkerNameProviderparameter.df.generateDataClasses(..),DataFrameSchema.generateInterfaces/generateDataClasses(..), and new parametersextensionProperties,visibility,useFqNames,nameNormalizer. Old no-arg overloads areHIDDEN.generateCode→generateInterfacesrow omits thatfieldswas removed and thatextensionPropertiesnow defaults tofalse(it wastrueingenerateCode)CodeWithConverter→CodeWithTypeCastGenerator(converter→typeCastGenerator: TypeCastGenerator,with()→declarationsWithCastExpression()).EMPTY_CONVERTERis deprecated in favor ofTypeCastGenerator.Empty.val `value`→val value).@Importannotation (@Import DataFrame.read(..)) was removed without deprecation.df.requireColumn { "name"<Type>() }: a runtime check that also adds the column to the compile-time schema. Use it to migrate gradually from the String API.df.properties(): ColumnsScope<T>, which shows only generated column properties in completion. Generated accessors now useColumnsScopeas the receiver (binary change for precompiled accessors).@DisableInterpretationcan now target functions, classes, and properties (for exampleinline fun <reified T> DataFrame<T>.foo()helpers).select("a"),remove("a"),groupBy("a"),rename, statistics, …) are interpreted by the plugin (@StringApiInterpretable), so untyped frames get refined compile-time schemas.split,gather,update,implode,fill*/drop*,convert(all),move/insert(incl.before,insideGroup),join, statistics on DataFrame/GroupBy,cumSum,distinct,parse,explode,allExcept,DataColumn.map*,valueCounts,toDataFrame/AddDsl/dataFrameOfbuilders, Java records, and more.IllegalStateException(handleExtensionPropertyException) instead of a bareClassCastException/NPE.DataFrame<T>?property as a value column, not aFrameColumn.3. Access API (KProperty / column accessors) deprecation
The guide has no section for this yet.
@AccessApiOverloadoverloads that takeKProperty,ColumnReference, orColumnAccessorare deprecated atWARNING(DEPRECATED_ACCESS_API, kept across releases). This covers every operation, the CS DSL (col/colGroup/cols/all*/…),column(property), androw[colRef].row[colRef]should becomecol[row],row["name"], orrow.name. Migrate to the String API or extension properties.ColumnReferencehelperswithValues(..),gt,lt,eq,neq,length(),lowercase(),uppercase()areERROR.ColumnReference<C>.map(infer) { }was made internal without anERRORphase.columnGroup.map { }now resolves toDataFrame.mapand returns aList. UseasDataColumn().map { }instead.4. DataFrame / column creation
Iterable<T>.toDataFrame()creates a singlevaluecolumn for value-like types (primitives,String, enums, date-time, arrays, classes without properties). The typed overloads (Iterable<Int>.toDataFrame(): DataFrame<ValueProperty<Int>>, …) were removed, and the result isDataFrame<T>.Mapvalues are value types:toDataFrame()/unfoldno longer expand them intosize/entries/… columns.unfolduses thetoDataFrame()rules.toDataFrame()/unfoldnow substitute generic type arguments (Pair<Int, String>→first: Int,second: String, notAny).excludeClassesworks in nested levels, and empty nested classes no longer produce empty groups.records are supported bytoDataFrame().toDataFrame { }(CreateDataFrameDsl) no longer extendsTraversePropertiesDsl. Top-levelexclude()/preserve()areERRORno-ops, so move them intoproperties { }. A new@CreateDataFrameDslMarkerforbids calling outer-DSL functions insideproperties { }.ValueColumn<DataFrame>is disallowed: creating one throws, andtoDataFrame { preserve(DataFrame::class) }throws. Slicing aValueColumn<AnyFrame?>without nulls yields aFrameColumn.ValueColumn.get(indices/range)now returnsDataColumn<T>.emptyDataFrame<T>()isinline reifiedand returns a 0-row frame withT's columns (=DataFrame.emptyOf<T>()). Before, it had no columns.DataColumn.empty(name)now returnsDataColumn<Nothing>(typeNothing, previouslyUnit). NewDataColumn.emptyOf<T>(name).df concat otherDf(different type parameters) now returnsDataFrame<Any>. Use thevararg/Iterableoverloads or.cast<T>().dataFrameOf("a" to col1, "g" to columnOf("x" to col2))andcolumnOf(vararg Pair<String, AnyBaseCol>), for hierarchical frames inline.Iterable<Map<String, Any?>>.toDataFrame()(about 10× faster than manual) andMap<String, Any?>.toDataRow().List<List<T>>.toDataFrame(header, containsColumns)inapi. Theioversion is deprecated (WARNING).df.chunked(startIndices, name),toSequence()/toSequenceOf<T>(), andunfold(vararg roots, maxDepth).chunked)DynamicDataFrameBuilder(checkDuplicateValues = true)no longer adds a renamed duplicate of an identical column. Newget(name)andadd(values, name).dataFrameOf(..).withValues(..)with the wrong number of values throws aDataFrameError.5. Columns Selection DSL
exceptNewbecame the stableexcept:colGroup.except { },"g".except("a"). It is no longerinfix, and@ExperimentalExceptCsDslwas removed. It keeps the group structure. The old accessor/ColumnsResolveroverloads areERRORand point to the{ }form.startsWith/endsWith(alreadyERRORin 0.15) were removed. UsenameStartsWith/nameEndsWithorcolsName*.TransformableColumnSet/TransformableSingleColumn(left over fromrecursively()) are internal.all,cols,colsOf,first,single, … now return plainColumnSet/SingleColumn(binary change).ColumnsResolver.isSingleColumn()/isSingleColumnWithGroup()were removed without deprecation.select { col(0) named "a" }/into "a".ColumnsContainer.mapToColumn(name, type: KType, …)overload became internal. A newDataFrame.mapToColumn(..)keepsdf.mapToColumnworking.mapToColumn→expr)colsAtAnyDepth(),colsInGroups(),single(),singleCol().6. Modification & reshaping operations
renameToCamelCase()was rewritten. It splits on any non-alphanumeric character and on case and digit boundaries, then lowercases words:ORDER_DATE→orderDate(wasoRDERDATE).DELIMITERS_REGEX/DELIMITED_STRING_REGEXwere removed, andColumnReference.renameToCamelCase()is new.merge { }.by(..)/asStrings()return the newMergeWithTransform, andnotNull()narrows types.move { }.before { },insert(col).before { }, andinsideGroupoverloads (toStart(insideGroup),toEnd(..),to(index, ..),moveTo*(insideGroup)), plusColumnsWithDifferentParentException.toLeft/toRightrenames)insert(col).under { path }now creates missing parent groups.under(path)deprecation)InsertException(NotAColumnGroupInsertException,DuplicateColumnPathInsertException).GroupBy.values()explains key/group column clashes.ungroupthrowsUngroupWrongColumnKindExceptionfor non-group columns.explodethrowsExplodeWrongColumnKindExceptionfor columns that are notListor frame columns (before, both silently did nothing).explode(verify = false)restores the lenient behavior.split { }.by(..).into(names)always creates all named columns and fills missing values withnull/default.update { }.with { }(andfill*) are wrapped inIllegalStateException("Could not update column '<name>'").append:nullinto aFrameColumnadds an empty frame,append()with no values returns the receiver, and a wrong value count throwsIllegalArgumentException(wasIllegalStateException).Gather,Corr, andSplitWithTransformare no longerdataclasses (copy/componentN/equalsare gone).ColumnMatch(a match binjoin) changed from a class to an interface, with aColumnMatch(left, right)factory function.JoinType.allowLeftNulls/allowRightNullsextensions became internal without deprecation.gather { listCols }.explodeLists(), so nocastis needed.DataRow.explode(..): named parameterselector→columns.AddDsl("col" from { },expr { } into, …) receivesAddExpression, soprev()?.newValue()works.DataRow<T>.transpose()→AnyRow.transpose()(the type parameter was dropped).df.filterNotNull { }was added as anERRORpointer todropNulls { }(for discoverability).col.reverse()insidesortBy { }is nowERRORand points toreversed().desc()→reversed())unfold { }has newroots/maxDepthparameters (binary change).7. GroupBy / Pivot / aggregation
groupBy.concatWithKeys(), which keeps the key columns.groupBy.countDistinct()/countDistinct { cols }.pivotCountsinsideaggregate { }now counts. Before, it returnedpivotMatchesbooleans.pivot/pivotCounts/pivotMatchesshortcuts take aPivotDslselector, sothenworks.df.asGroupBy { }moves the group column to the end, so aggregation results have a predictable column order.PivotGroupBy.maxOf { }/minOf { }producenullfor empty groups instead of throwing.GroupBy<*, *>.cast<T, G>().8. Statistics
The guide covers only
BigDecimal/BigIntegerand therow*Ofrenames.skipNA→skipNaNinsum/mean/std(allFor/Of/rowvariants on column, frame,GroupBy,Pivot), so named callsmean(skipNA = true)no longer compile.skipNaNwas added tomin/max/median/percentile.cumSumkeepsskipNA.medianof number columns returns an interpolatedDouble(quantile R8). In 0.15 it returned the column type with integer division. OtherComparablecolumns return the lower middle value (R3). NewmedianBy/medianByOrNull.percentile(p)family (percentile*,percentileFor/Of/By,rowPercentileOf), with the same typing asmedian.sumresult types:Byte/Short→Int, mixedNumber→ unified type, empty input → typed0.meanalways returnsDouble(NaNfor empty input), andDataColumn.meanOrNull()was removed.stdusesddof(default1).Numbercolumns are converted to the smallest common type (Int+Float→Double), which fixesrowSum.sum()/mean()/std()without columns select only primitive or mixed-number columns.min/max/minBy/maxOf/…: bound isT : Comparable<T & Any>?, many areinline reified, and mixed-number columns are no longer supported.rowMin/rowMaxrename)minFor/maxFor/medianFor/percentileForaccept columns of different comparable types (Comparable<*>?).rowXOf<T>()selects row values by runtime value type, not by column type.describe()has newp25/p75columns (ColumnDescription.p25/p75) aroundmedian.cumSumfollows thesumtyping:Byte/Short→Int, and nullableDouble/Float→ non-null withNaN.AnyCol.isBigNumber()was removed without deprecation. NewisPrimitiveNumber(),isMixedNumber(),isPrimitiveOrMixedNumber()....dataframe.math(Iterable.mean(),std,median,varianceAndMean,BasicStats) were made internal without deprecation.digitizewas removed (all overloads) without deprecation.median/percentile/cumSum: see errata.9. Types, parsing & conversion
parseExperimentalUuidnow defaults totrue:parse(),convertTo, andreadCsv/readTsv/… turn UUID-formatted strings intokotlin.uuid.Uuid. Opt out withParserOptions(parseExperimentalUuid = false)orDataFrame.parser.parseExperimentalUuid = false. The option and UUID parsing itself are new since 0.15.useFastDoubleParsernow defaults totrue, with a fallback that also tries alternative decimal symbols for the locale (more lenient).Charcolumns:df.parse()parses them ('1'→Int,'T'→Boolean),convert { }.to<>()/convertTo<>()fall back toString(for example to enums), andDataColumn<Char?>.parse()/tryParse()are new.convertshortcuts:toJavaLocalDate/LocalTime/LocalDateTime/Instant/Duration(withpattern/DateTimeFormatter+locale),toDuration,toYearMonth,toUtcOffset,toDateTimeComponents(..),toLocalDate(format: DateTimeFormat), …convertTo<T>()/convert.to<T>()acceptparserOptions.toLocalDate(pattern)/toJavaLocalDate)locale(9 functions) andDataFrame.parser.addDateTimePattern(..)are alreadyERROR.ParserOptions.dateTimeFormatter/dateTimePatternproperties were removed (not only the constructor parameters), andcopy()changed. New exceptions:ColumnTypeMismatchesColumnValuesExceptionandTypeConversionException.extraInformation.convertshortcut overloads were merged into nullable-receiver variants, and the non-nullConvert<T, Any>.toInt()/toStr()/…overloads were removed (source-compatible, binary-incompatible).URI(str).toURL(), so malformed URLs (for example with spaces) now throw.10. Schema, cast & comparison
DataFrameSchema.equals()takes column order into account (also for nested groups and frames). Useschema.compare(other).matches()for an order-insensitive comparison.Equals→Matches)compare(other, strictlyEqualNestedSchemas)→compare(other, comparisonMode = ComparisonMode.LENIENT)(LENIENT/STRICT/STRICT_FOR_NESTED_SCHEMAS).ColumnSchemaissealed, andCompareResult.plusis new.compileTimeSchema()is ordered likeschema()and has a neworderedparameter.changeType(type)is a public member ofValueColumn(it was on the internalDataColumnInternal).11. IO: general
ERROR:DataFrame.read(file/url/path/String),DataRow.read(..),File/URL/Path.readDataFrame()/readDataRow(), the wholeSupportedDataFrameFormatservice (CsvDeephaven,TsvDeephaven,Excel,ArrowFeather,Jdbc,JSONread overrides),buildCodeForDB, andCodeGenerator.urlDfReader. UsereadCsv/readJson/readExcel/readParquet/readSqlTable/…read(delimiter = ..); see errata)ColumnNameGenerator(a,a1,a2).readExcelno longer throwsDuplicateColumnNamesException, and JDBC givesid1instead ofid_1. The deprecatedreadExcel(nameRepairStrategy = ..)overloads ignore the argument, so just drop it.java.nio.file.Pathoverloads for all IO (read,readJson/writeJson,readExcel/writeExcel, Arrow, Geo, CSV,importDataSchema).SupportedDataFrameFormats must implementreadDataFrame(path: Path, ..)(theFilevariant now has a default).isOpenApi(File)→isOpenApi(Path).12. IO: CSV/TSV
readCsv/readTsv/readDelim:compressionmoved to the end, and a newcharset: Charset?comes afterheader(the BOM is auto-detected, UTF-8 fallback). Old overloads areHIDDEN, so positional calls break and named ones keep working.Instantinstead of a shiftedLocalDateTime).Compressionmoved fromdataframe-csvtocore(same package) and became afun interface(Gzip/Zip/Noneobjects,Compression.of(..)).Compression.Custom(..)was removed, so use the SAM form.readDelim(InputStream/Reader)andCSVTypeareWARNING. The oldreadDelimStrisHIDDEN, so calls resolve to the Deephaven one.InputStream.skippingBomCharacters()also strips UTF-16/32 BOMs.13. IO: JSON
unifyNumbers = trueon allreadJson*and onJSON(..): mixed numbers get a common type instead ofNumber, andFloatcan be inferred. Inferred types may change.readJson(..):nullwhere an array is expected givesnull(List<T>?) instead of an empty list.toJsonwritesnulllists asnull, and a user column namedvalue/arrayno longer breaks writing.readJson(keyValuePaths = ..)producesname/valuecolumns (waskey/value).@JsonOptions(typeClashTactic = ..)takes theStringconstantsJsonOptions.TypeClashTactics.ARRAY_AND_VALUE_COLUMNS/ANY_COLUMNSinstead of the enum.readJson(stream, .., format: Json)for a custom kotlinx-serializationJson(lenient parsing, comments, …).14. IO: Excel
readExcel(.., parseEmptyAsNull = true): empty-string cells becomenull, which changes inferred types and nullability.writeExcelto a new file uses the streamingSXSSFWorkbook(lower memory).keepFile = truewith a missing or empty file creates a workbook instead of failing.15. IO: Arrow / Parquet
DataFrame.readParquet(urls/paths/files, nullability, batchSize)(Arrow Dataset, native).ARROW_PARQUET_DEFAULT_BATCH_SIZEis public.StructVector↔ColumnGroup: structs are read as column groups (before, value columns of map-like values), and groups are written as structs. Nullable structs respect the null bit (children become nullable).ListVector/LargeListVectorof primitives are read asList<T>, and lists of structs asFrameColumn.kotlin.time.Instant(before,NotImplementedError).Instantcolumns are written asTimestamp(MICROSECOND, "UTC")instead ofUtf8strings, and Timestamp targets are pinned to UTC.ConvertingMismatch.PrecisionReduced/ValueOutOfRangesubclasses, so an exhaustivewhenover the sealed class is source-incompatible.16. IO: JDBC / SQL databases
The guide section is a TODO, and almost none of this went through a deprecation cycle.
DataFrame.getSchemaForSqlTable/ForSqlQuery/ForResultSet/ForAllSqlTables(..)→DataFrameSchema.readSqlTable/readSqlQuery/readResultSet/readAllSqlTables(..), andConnection/DbConnectionConfig/ResultSet.getDataFrameSchema(..)→readDataFrameSchema(..).limit: Int = Int.MIN_VALUE→limit: Int? = nullin all readers. A new trailingconfigureStatement: (PreparedStatement) -> Unit. The old signatures were removed.javax.sql.DataSourceoverloads for every reader and schema reader (connection pools).TIMESTAMP→kotlin.time.Instant(wasjava.sql.Timestamp), JavaLocalDateTime→ kotlinxLocalDateTime,UUID→kotlin.uuid.Uuid, MySQL/MariaDBBIGINT UNSIGNED→BigInteger(BigInteger statistics are not supported), MariaDBINT UNSIGNED→Long. PostgreSQL extension types are supported.DbConnectionConfig(.., readOnly = true): connections from a config are read-only by default (auto-commit off, rolled back after reading). It is no longer adata class(componentNis gone), andtoString()masks the password.validation: SqlValidation = SqlValidation.NoneonreadSqlQuery/readDataFrame, so queries are not validated by default. In 0.15 they had to start withSELECTand contain no;.SqlValidation.ReadOnlyallows only oneSELECT/WITH/VALUES/TABLE/EXPLAINstatement. Table names inreadSqlTableare always validated.DbTypeoverhaul (customDbTypes must migrate):convertSqlTypeToColumnSchemaValue/convertSqlTypeToKType/extractValueFromResultSetwere removed in favor ofgetExpectedJdbcType→getValueFromResultSet→preprocessValue→getTargetColumnSchema/buildDataColumn.sqlQueryLimit→buildSqlQueryWithLimit. New hooks:quoteIdentifier,configureReadStatement,createConnection,defaultFetchSize,defaultQueryTimeout,tableTypes,getTableColumnsMetadata.AdvancedDbTypeand theJdbcToDataFrameConverterpipeline (DbResultSetReader,DbValuePreprocessor,DbColumnBuilder,jdbcToDfConverterFor,with*).TableColumnMetadata/TableMetadatamoved from...dataframe.ioto...dataframe.io.db.DuckDb(auto-detected fromjdbc:duckdb:): STRUCT → column group, LIST/ARRAY, MAP, JSON,GEOMETRY,TIME_NS,VARIANT, and read-only mode.H2(dialect: DbType = MySql)→H2(mode: H2.Mode = H2.Mode.Regular). The default is no longer the MySQL dialect, andH2(dialect)/H2.MODE_*are deprecated. URLs withoutMODE=are supported.Sqlitechanged from anobjectto a class (AdvancedDbType): useSqlite.defaultorSqlite(), plusSqlite.withCustomConverters { forType<..>(..) {}; forColumn<..>(..) {} }. Declared types are mapped:BOOLEAN→Boolean,DATE/DATETIME/TIME/TIMESTAMP→ kotlinx date-time/Instant.readAllSqlTableswithoutcataloguefalls back toconnection.catalog(MySQL/MariaDB/MSSQL: only the URL's database).tableTypesdefaults toTABLE,BASE TABLE.id,id1(wasid,id_1).17. IO: Geo & OpenAPI
dataframe-geo-jupyter.%use dataframe(enableExperimentalGeo=true)no longer loads Geo, so use%use dataframe-geo.dataframe-geono longer depends ondataframe-jupyter.GeoDataFrame.modify { }also receives the frame asit.readShapefile(dir)finds<dir>.shpinside a directory.writeGeoJsonwrites ordered feature IDs, so row order is preserved.@ImportDataSchema(enableExperimentalOpenApi = true), or%use dataframe(enableExperimentalOpenApi=true)in notebooks.AdditionalProperty<T>extendsNameValueProperty:key→name(and the generated column isname).18. Rendering: HTML, format, print, Markdown
print()/renderToString()/DataFrame.toString()output changed: column types are shown by default (columnTypes = true, on a separate row), an "N columns × M rows" footer appears on truncation, andDataColumn.toString()includes the type. String-snapshot tests will break.print/renderToString:rowsLimit/valueLimitareInt?(nullmeans unlimited), and there are newborderStyle: StringBorderStyleandrowIndexparameters.renderToStringis public. The title changed fromData Frame [..]toDataFrame [..].renderToMarkdown(..)andMarkdownAlignment.format { }works on nested columns and column groups.RowColFormatter(andDisplayConfiguration.cellFormatter) receivesColumnWithPath<C>instead ofDataColumn<C>.format { }.where { }.where { }now combine with AND (the last one used to win).textColoralso applies to numbers and nulls.FrameColumnframes with same-named columns. Format nested frames explicitly.formathelpers:at(rows),notNull()/notNull { },FormattedFrame.format(vararg String).formatHeader { cols }.with { }andDisplayConfiguration.headerFormatter. TheFormatClause/FormattedFrameconstructors changed.FormattedFrame.toHTML/toStandaloneHTMLwere removed without deprecation, so usetoHtml/toStandaloneHtml(now withcellRendererandgetFooter).DataFrame.toHTML)FormattingDSL/RGBColorare now typealiases (the old classes are gone from bytecode, so it is a binary change).FormattedFramerenders formatted in Kotlin Notebook (toJsonWithMetadata(isFormatted)).19. Kotlin Notebook / Jupyter
GroupByvariables (for examplegroupBy.keys.kin the next cell).ExperimentalTime/ExperimentalUuidApiforInstant/Uuidcolumns.ColumnGroupcast toDataFramerenders as a frame of its sub-columns.List/DataFramecolumns by size.20. Utilities & misc
DataFrame.none { }andDataColumn.none { }.exceptions.DataFrameError, implemented byColumnNotFoundException,DuplicateColumnNamesException, andColumnsWithDifferentParentException.KeyValueProperty→NameValueProperty: the generatedkeycolumn/property extensions are deprecated in favor ofname.NamedValue.copy(..)is no longer public.Notable things deliberately left out
rename(vararg Pair)swapped names), Fix accidental and too strict order check fromcast#1716 (castorder check), Fixed concatKeepingSchema turning ColumnGroups into DataColumns #1764 (concatKeepingSchema), and Fixed reading Parquet file on Windows #1381 (Parquet on Windows).groupBy35–50% faster), Lazy Statistics #1656 (lazy statistics cache), Optimize type inference in joinWith #1633 (joinWithtype inference), and Optimize Iterable<Map>.toDataFrame conversion #1635 (Iterable<Map>.toDataFrame10× faster).strictValidation(replaced bySqlValidation),rowPercentile, the Removed redundant move overloads where insideGroup argument doesn't affect the result #1793insideGroupvararg Stringoverloads, MergingKeyValuePropertyandNameValuePair#1532/Revert "MergingKeyValuePropertyandNameValuePair" #1544 (reverted), andSqlite.withCustomTypes.@Converter,@StringApiInterpretable,UnresolvedColumnsPolicy, the publicColumnResolutionContext, anddataframe-compiler-plugin-core.*Grammar*/DocumentationUrls*interfaces were removed from the API dump ((Nested) typealias KDocs #1543, KoDEx bump #1529).AnyFrame,AnyRow,AnyBaseCol; InlineAnyFrame,AnyRow, andAnyBaseColtype aliases in public API #1863, Inline ColumnFilter typealias in public API function signatures #2003): identical types, so there is no impact.AggregateGroupedDslwith compiler plugin #1454, Type safe col ref does not work in rename/remove when switched to compiler plugin #1461, Provide clearer error message for unsupported local@DataSchema data class#1721, Error when casting dataframe to a @DataSchema class located in another package #1877, …): they ship with Kotlin releases. The Compiler plugin section of the guide could link to them.🤖 Generated with Claude Code