

Objectivity, an established Silicon Valley firm with experience in high-performance distributed object-oriented databases, today debuted a new Hadoop-based product that addresses one of the looming challenges in the Internet of Things (IoT): How to handle metadata management of big and fast streaming data.
It may not seem obvious from the outside, but one of the challenges in tackling big streaming data is metadata management. Much of the unstructured and semi-structured data that flows (or will flow) across the IoT is not readily usable in its raw form. Time-series data, in particular, often needs to be transformed before it can be consumed, by analytic applications or otherwise.
Marking up one data stream wouldn’t be so bad. The associated metadata can be cataloged and stored without too much trouble. But as organizations mix and match multiple streams, the whole pipeline threatens to become a messy quagmire.
That’s roughly the challenge that Objectivity hopes to address with ThingSpan, the new YARN-certified Hadoop application that the company unveiled this morning. Inspired by Objectivity’s success in the object-oriented database (Objectivity/DB) and graph database (InfiniteGraph) spaces–and borrowing technology from those two products—the new ThingSpan product should help keep customers’ IoT and streaming data projects on the straight and narrow, says Jin Kim, vice president of marketing and partner development for Objectivity.
“A lot of this senor data comes in time-series form and time-series data has high dimensionality that’s not well suited for many of the analytic algorithms. So one of the key aspects is dimensionality reduction,” Kim tells Datanami.
“But when we do that kind of dimensionality reduction, we need to create a lot of metadata, because some of the analytics is run on the metadata and not the actual raw data,” Kim continues. “What we’re trying to do is to basically create the frameworks so we can enrich it with semantics technology so it can be more oncology driven. It’s about MDM and the automatic creation of metadata and the data model that is necessary for complex fusion processes.”
ThingSpan will be that Hadoop-based repository of metadata created from that streaming data. The company won’t do any of the actual analytics—it will leave that up to the individual customer, who typically have strong preferences. “We kind of want to keep it analytics agnostic so they can bring their favorite buffet of analytic tools with them,” Kim says. “We’re not about to tell our customer that we have a better set of enrichment or clustering or anomaly detection techniques than they do.”
With that said, the software is being developed to work with the graph analytic and machine learning tools available in Apache Spark. It’s also being developed to work with Apache Kafka, as well as Project Apex, a streaming analytic application developed by DataTorrent.
Objectivity thinks it can offer companies that are building streaming analytics and IoT applications a better and more scalable MDM framework that what is currently available, which tend to be mostly modified NoSQL databases, Kim says. Objectivity has been solving these sorts of problems for customers in the intelligence and military sector for the past decade, and now sees an opportunity as real-time analytic applications become more common in the commercial market.
“Objectivity has been dealing with the domain of how to integrate and fuse fast time-series data form sensor networks and enrich them with contextual information for a long time,” Kim says. “It’s been doing this on beyond petabyte-scale data, approaching data ingestion rates well over 1 billion events per second.”
The popularity of object-based technologies has come and gone over the years, but the folks at Objectivity see the technology now being used to bring performance and scalability advantages to the burgeoning field of IoT and streaming analytic applications.
“One of our intelligence customers told us recently that for every piece of data they ingest, they generate six separate metadata items for all the relationships they need to maintain,” Kim says. That’s why “object-based technology is coming into vogue again, [because] as people introduce concepts like data lakes and the idea of ingesting multi-various types of data…you are beginning to deal with much more complex metadata.”
ThingSpan has been certified to run on the Hadoop distributions from Hortonworks and Cloudera, and Objectivity is working with MapR Technologies, Kim says. This will jump start the MDM efforts of companies that are developing IoT and streaming analytic applications on Hadoop, without subjecting them to the steep learning curve that organizations in the intelligence and oil and gas fields had to deal with, and without incurring the high price tags that accompany enterprise streaming analytic products from big-name vendors like IBM and Software AG, Kim says.
“The industry needs a standard stack for running advanced and streaming analytics,” says Kim, who worked previously worked at Skytree. “Intel‘s Trust Analytics initiative is a good [start]. But we need more standards…To effectively run complex analytics, you need to automatically generate and maintain the complex metadata and relationships. We think more and more that metadata structure will be ontology-driven. It has to be as the data set gets richer and just from a provenance point of view. You have to do it.”
ThingSpan will become generally available in October. The company will be showcasing the product next week at the Strata + Hadoop World conference in New York City.
Related Items:
One Deceptively Simple Secret for Data Lake Success
What’s Driving the Rise of Real-Time Analytics
Unstructured Data Analytics Shouldn’t Be Such a Mess
June 16, 2025
- Linux Foundation Announces 2025 Open Source Summit Europe Schedule for Amsterdam
- Bloomberg Integrates Natural Language Search Across Terminal Research Content
June 13, 2025
- PuppyGraph Announces New Native Integration to Support Databricks’ Managed Iceberg Tables
- Striim Announces Neon Serverless Postgres Support
- AMD Advances Open AI Vision with New GPUs, Developer Cloud and Ecosystem Growth
- Databricks Launches Agent Bricks: A New Approach to Building AI Agents
- Basecamp Research Identifies Over 1M New Species to Power Generative Biology
- Informatica Expands Partnership with Databricks as Launch Partner for Managed Iceberg Tables and OLTP Database
- Thales Launches File Activity Monitoring to Strengthen Real-Time Visibility and Control Over Unstructured Data
- Sumo Logic’s New Report Reveals Security Leaders Are Prioritizing AI in New Solutions
June 12, 2025
- Databricks Expands Google Cloud Partnership to Offer Native Access to Gemini AI Models
- Zilliz Releases Milvus 2.6 with Tiered Storage and Int8 Compression to Cut Vector Search Costs
- Databricks and Microsoft Extend Strategic Partnership for Azure Databricks
- ThoughtSpot Unveils DataSpot to Accelerate Agentic Analytics for Every Databricks Customer
- Databricks Eliminates Table Format Lock-in and Adds Capabilities for Business Users with Unity Catalog Advancements
- OpsGuru Signs Strategic Collaboration Agreement with AWS and Expands Services to US
- Databricks Unveils Databricks One: A New Way to Bring AI to Every Corner of the Business
- MinIO Expands Partner Program to Meet AIStor Demand
- Databricks Donates Declarative Pipelines to Apache Spark Open Source Project
June 11, 2025
- What Are Reasoning Models and Why You Should Care
- The GDPR: An Artificial Intelligence Killer?
- Fine-Tuning LLM Performance: How Knowledge Graphs Can Help Avoid Missteps
- It’s Snowflake Vs. Databricks in Dueling Big Data Conferences
- Snowflake Widens Analytics and AI Reach at Summit 25
- Inside the Chargeback System That Made Harvard’s Storage Sustainable
- Top-Down or Bottom-Up Data Model Design: Which is Best?
- Why Snowflake Bought Crunchy Data
- Change to Apache Iceberg Could Streamline Queries, Open Data
- Stream Processing at the Edge: Why Embracing Failure is the Winning Strategy
- More Features…
- Mathematica Helps Crack Zodiac Killer’s Code
- It’s Official: Informatica Agrees to Be Bought by Salesforce for $8 Billion
- AI Agents To Drive Scientific Discovery Within a Year, Altman Predicts
- Solidigm Celebrates World’s Largest SSD with ‘122 Day’
- DuckLake Makes a Splash in the Lakehouse Stack – But Can It Break Through?
- The Top Five Data Labeling Firms According to Everest Group
- Who Is AI Inference Pipeline Builder Chalk?
- ‘The Relational Model Always Wins,’ RelationalAI CEO Says
- IBM to Buy DataStax for Database, GenAI Capabilities
- VAST Says It’s Built an Operating System for AI
- More News In Brief…
- Astronomer Unveils New Capabilities in Astro to Streamline Enterprise Data Orchestration
- Yandex Releases World’s Largest Event Dataset for Advancing Recommender Systems
- Astronomer Introduces Astro Observe to Provide Unified Full-Stack Data Orchestration and Observability
- BigID Reports Majority of Enterprises Lack AI Risk Visibility in 2025
- Databricks Unveils Databricks One: A New Way to Bring AI to Every Corner of the Business
- MariaDB Expands Enterprise Platform with Galera Cluster Acquisition
- Snowflake Openflow Unlocks Full Data Interoperability, Accelerating Data Movement for AI Innovation
- Gartner Predicts 40% of Generative AI Solutions Will Be Multimodal By 2027
- Databricks Announces 2025 Data + AI Summit Keynote Lineup and Data Intelligence Programming
- Databricks Announces Data Intelligence Platform for Communications
- More This Just In…