从Dataflow api在数据存储区中保存长度超过1500字节的字符串时出错

问题描述 投票:2回答:2

当我尝试保存一个非常长的字符串时,Dataflow作业抛出此错误消息:属性“myProperty”的值超过1500字节。,code = INVALID_ARGUMENT。

在关注Google的qazxsw poi示例并保存超过1500字节的字符串时出错。

我知道在使用Datastore API时,我可以通过将属性保存为DatastoreWordCount来保存长度超过1500字节的字符串。但是,在com.google.appengine.api.datastore.Text样本或DatastoreWordCount类文档中没有其他选择可以表明支持DatastoreHelper类型。

可能是一种使用该API保存这么长的字符串的方法,以便它可以被读作Text

完整的错误消息如下:

com.google.appengine.api.datastore.Text
google-app-engine google-cloud-datastore google-cloud-dataflow
2个回答
4
投票

您可以通过从索引中排除值来保存长度超过1500字节的字符串:

java.lang.RuntimeException: com.google.cloud.dataflow.sdk.util.UserCodeException: java.lang.RuntimeException: com.google.cloud.dataflow.sdk.util.UserCodeException: java.lang.RuntimeException: com.google.cloud.dataflow.sdk.util.UserCodeException: java.lang.RuntimeException: com.google.cloud.dataflow.sdk.util.UserCodeException: com.google.datastore.v1.client.DatastoreException: The value of property "dalekTestExecutions" is longer than 1500 bytes., code=INVALID_ARGUMENT
    at com.google.cloud.dataflow.sdk.runners.worker.SimpleParDoFn$1.output(SimpleParDoFn.java:162)
    at com.google.cloud.dataflow.sdk.util.DoFnRunnerBase$DoFnContext.outputWindowedValue(DoFnRunnerBase.java:288)
    at com.google.cloud.dataflow.sdk.util.DoFnRunnerBase$DoFnContext.outputWindowedValue(DoFnRunnerBase.java:284)
    at com.google.cloud.dataflow.sdk.util.DoFnRunnerBase$DoFnProcessContext$1.outputWindowedValue(DoFnRunnerBase.java:508)
    at com.google.cloud.dataflow.sdk.util.GroupAlsoByWindowsAndCombineDoFn.closeWindow(GroupAlsoByWindowsAndCombineDoFn.java:205)
    at com.google.cloud.dataflow.sdk.util.GroupAlsoByWindowsAndCombineDoFn.processElement(GroupAlsoByWindowsAndCombineDoFn.java:192)
    at com.google.cloud.dataflow.sdk.util.SimpleDoFnRunner.invokeProcessElement(SimpleDoFnRunner.java:49)
    at com.google.cloud.dataflow.sdk.util.DoFnRunnerBase.processElement(DoFnRunnerBase.java:139)
    at com.google.cloud.dataflow.sdk.runners.worker.SimpleParDoFn.processElement(SimpleParDoFn.java:190)
    at com.google.cloud.dataflow.sdk.runners.worker.ForwardingParDoFn.processElement(ForwardingParDoFn.java:42)
    at com.google.cloud.dataflow.sdk.runners.worker.DataflowWorkerLoggingParDoFn.processElement(DataflowWorkerLoggingParDoFn.java:47)
    at com.google.cloud.dataflow.sdk.util.common.worker.ParDoOperation.process(ParDoOperation.java:55)
    at com.google.cloud.dataflow.sdk.util.common.worker.OutputReceiver.process(OutputReceiver.java:52)
    at com.google.cloud.dataflow.sdk.util.common.worker.ReadOperation.runReadLoop(ReadOperation.java:224)
    at com.google.cloud.dataflow.sdk.util.common.worker.ReadOperation.start(ReadOperation.java:185)
    at com.google.cloud.dataflow.sdk.util.common.worker.MapTaskExecutor.execute(MapTaskExecutor.java:72)
    at com.google.cloud.dataflow.sdk.runners.worker.DataflowWorker.executeWork(DataflowWorker.java:287)
    at com.google.cloud.dataflow.sdk.runners.worker.DataflowWorker.doWork(DataflowWorker.java:223)
    at com.google.cloud.dataflow.sdk.runners.worker.DataflowWorker.getAndPerformWork(DataflowWorker.java:173)
    at com.google.cloud.dataflow.sdk.runners.worker.DataflowWorkerHarness$WorkerThread.doWork(DataflowWorkerHarness.java:193)
    at com.google.cloud.dataflow.sdk.runners.worker.DataflowWorkerHarness$WorkerThread.call(DataflowWorkerHarness.java:173)
    at com.google.cloud.dataflow.sdk.runners.worker.DataflowWorkerHarness$WorkerThread.call(DataflowWorkerHarness.java:160)
    at java.util.concurrent.FutureTask.run(FutureTask.java:266)
    at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)
    at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)
    at java.lang.Thread.run(Thread.java:745)

如果您需要与App Engine的Value longString = Value.newBuilder() .setStringValue(...) .setExcludeFromIndexes(true) .build(); 类型兼容,您还需要将含义设置为15:

com.google.appengine.api.datastore.Text

2
投票

DataStore为每个属性创建索引,因此属性上的默认限制为1500字节。现在,如果您需要存储类似大JSON的数据,那么您可以通过以下方式指定此属性不需要索引:

Value longString = Value.newBuilder()
    .setStringValue(...)
    .setExcludeFromIndexes(true)
    .setMeaning(15)
    .build();

这样您就可以保存更大的数据而不是默认的1500字节限制。


0
投票

确切地说:

Entity newEntity =
                Entity.newBuilder(key)
                        .set("time", Timestamp.parseTimestamp("1970-01-01T00:00:00Z"))
                        .set("message", StringValue.newBuilder(JSON).setExcludeFromIndexes(true).build())
                        .build();
© www.soinside.com 2019 - 2024. All rights reserved.