Data Integration

 Integrate data from multiple enterprise sources into a unified graph model for clearer visualization, analysis, and decision-making.  

What is data integration?

Data integration is the process of bringing data from separate systems into a form where it can be used together. In practice, this means connecting to databases, files, APIs, and other enterprise sources through purpose-built connectors, then preparing the data for analysis, visualization, or use within an application. Depending on the approach, this can involve techniques such as ETL (Extract, Transform, Load), data replication, or virtualization to consolidate data from any source.

Enterprise data lives across many connected systems, and its value often lies in the relationships among records, systems, and entities. By bringing data into a unified graph model, Tom Sawyer Perspectives helps organizations work directly with those relationships rather than treating each source as a separate silo.

Data connectors bring data from enterprise sources into a unified in-memory graph model.

Data connectors bring data from enterprise sources into a unified in-memory graph model.

Considerations when choosing a data integration solution

 Selecting the right data integration solution requires balancing business requirements, technical capabilities, performance, scalability, security, and long-term cost. The right solution should support your existing data sources and integration needs while remaining flexible as those requirements evolve. 

Consider these points as you evaluate data integration solutions.

Business Objectives and Requirements

Business Objectives and Requirements

Identify what you aim to achieve with the integration, such as improved data quality, real-time analytics, or streamlined operations.

Consider the types of data you need to integrate such as, customer data or financial data, and from which sources (e.g., CRM, ERP).

Data Volume and Complexity

Data Volume and Complexity

Assess the volume of data that will be handled and the system’s capability to scale as data grows.

Consider the complexity of the data structures and the need for data transformation.

Integration Capabilities

Integration Capabilities

Check if the solution supports different integration styles such as ETL (Extract, Transform, Load), real-time streaming, or batch processing.

Evaluate the ability to connect to various data sources and destinations, including cloud services, databases, and third-party APIs.

Performance and Scalability

Performance and Scalability

Look at the performance benchmarks, particularly how the system performs under load.

Ensure the solution can scale horizontally or vertically based on future needs.

Compliance and Security

Compliance and Security

Determine the security measures provided, including data encryption and secure data transfer protocols.

Assess data governance capabilities, such as data auditing, lineage, and cataloging.

Ease of Use and Maintenance

Ease of Use and Maintenance

Consider the user interface and ease of use for technical and non-technical users.

Evaluate the maintenance support offered by the provider, including customer service, updates, and patches.

Cost

Cost

Review the pricing structure, including initial setup costs, licensing fees, and ongoing operational costs.

Consider the total cost of ownership over time, including upgrades and additional services.

Vendor Reputation and Support

Vendor Reputation and Support

Research the vendor’s reputation in the market, their stability, and customer reviews.

Look at the level of technical support provided, including the availability of training resources and community support.

Flexibility and Customization

Flexibility and Customization

Check if the solution can be customized to fit specific business needs and how easily these customizations can be implemented.

Assess the flexibility of the solution to adapt to new technologies and integration patterns in the future.

Trial and Testing

Trial and Testing

If possible, conduct a proof of concept or trial to see how the solution fits with your existing systems and meets your integration needs.

Data integration with Tom Sawyer Perspectives

Tom Sawyer Software has spent more than three decades working with data and has a proven track record of providing solutions that address the data silo issue.

Tom Sawyer Perspectives supports multiple data sources and formats, providing the ability to integrate data from a wide range of sources, from graph and relational databases to cloud services, APIs, and more.

Tom Sawyer Perspectives can handle large volumes of data efficiently and scale as your business needs grow.

Once data is integrated, users gain real-time access to a unified view of the data and can use powerful visualization and analysis capabilities to reveal valuable insights.

From data integration to visualization with Tom Sawyer Perspectives.

From data integration to visualization with Tom Sawyer Perspectives.

10

relational databases

10

graph databases

7

data formats

Supported data integrators

 The Tom Sawyer Perspectives platform is data-agnostic. It includes a comprehensive set of data integrators that can populate a model from a data source. Some integrators provide a bidirectional bridge between a data source and a model and support writing model data back to the data source. 

 

The Perspectives data-agnostic platform supports dozens of data integration sources.

Tom Sawyer Perspectives integrators can access data in one or more of these sources:

  • Graph databases, such as Neo4j, Amazon Neptune and Neptune Analytics openCypher, and Kuzu

  • Graph databases supported by the Apache TinkerPop framework, such as Amazon Neptune Gremlin, Microsoft Azure Cosmos DB, JanusGraph, OrientDB, and TinkerGraph

  • Microsoft Excel

  • MongoDB databases

  • JSON files

  • RDF sources, such as Oracle, Stardog, AllegroGraph, QLever, and MarkLogic databases and RDF-formatted files

  • RESTful API servers

  • SQL JDBC-compliant databases, including Microsoft SQL Server, MySQL, Oracle, and Postgres

  • Structured text documents

  • XML

How data integration works with Tom Sawyer Perspectives

Tom Sawyer Perspectives makes it easy to integrate data from relational databases, graph databases, and RESTful APIs into a unified model.

For relational databases

For relational databases, the data integration process follows three core steps, which you can repeat as needed for additional data sources.

1

Connect to your data source

Use the provided integrators to connect to your data source.

2

Extract the schema

Use automatic schema extraction to read the structure of your data and pull out the metadata information automatically. Use the schema editor to view and manually adjust the Tom Sawyer Perspectives schema to match your application needs.

3

Bind the data

Bind the data source and the schema. This process determines how many model elements of each model element type are created in the model, and determines the values of all attributes in the model.

4

Repeat

Repeat the steps to connect to as many data sources as you need.

For graph databases

For graph databases, Tom Sawyer Perspectives supports automatic binding by default. This allows you to use a query to preview database content and visualize the data in a drawing view without manually creating a schema or defining data bindings.

Tom Sawyer Perspectives' dynamic data integration capabilities let you integrate data in about 20 seconds. From there, you can move from a Cypher- or Gremlin-compatible database to a fully customized, interactive visualization application in about 100 seconds more.

 

Neo4j integrator

The Neo4j integrator populates a model from a Neo4j database using the Bolt protocol or Neo4j REST API. If you use the integrator with the Bolt protocol, you can automatically create a Tom Sawyer Perspectives schema where the data bindings work as follows:

  • Element bindings correspond to a read Cypher query.

  • Attribute bindings correspond to a column in the result set of the Cypher query.

A Cypher query can use Cypher parameters.

Neptune Gremlin integrator

The Neptune Gremlin integrator populates a model from an Amazon Neptune database that supports a Property Graph model. The Neptune integrator automatically creates a Tom Sawyer Perspectives schema where the data bindings work as follows:

  • Element bindings correspond to a Gremlin query.

  • Attribute bindings correspond to a column in the result set of the Gremlin query.

After you define the data bindings, you can configure the integrator to write changes back to the database.

Neptune openCypher integrator

The Neptune openCypher integrator uses the Bolt protocol to populate a model from an Amazon Neptune openCypher database that supports a Property Graph model. With the integrator, you can automatically create a Tom Sawyer Perspectives schema where the data bindings work as follows:

  • Element bindings correspond to a read Cypher query.

  • Attribute bindings correspond to a column in the result set of the Cypher query.

TinkerPop integrator

The TinkerPop integrator works with TinkerPop-compliant graph databases such as TinkerGraph, JanusGraph, and Neo4j-Gremlin. You should also use a TinkerPop integrator for OrientDB Server version 3.0 or earlier. The TinkerPop integrator can automatically create a Tom Sawyer Perspectives schema where the data bindings work as follows:

  • Element bindings correspond to a Gremlin query.

  • Attribute bindings correspond to a column in the result set of the Gremlin query.

A Gremlin query can use parameters.

Kuzu integrator

The Kuzu integrator populates a model from a local Kuzu database. The Kuzu integrator automatically creates a Tom Sawyer Perspectives schema where the data bindings work as follows:

  • Element bindings correspond to a read Cypher query.

  • Attribute bindings correspond to a column in the result set of the Cypher query.

A Cypher query can use parameters. After you define the data bindings, you can configure the integrator to write changes back to the database.

For RESTful API

Tom Sawyer Perspectives REST integrator is a service that populates a model from a RESTful API server. It supports responses in JSON and XML formats.

The REST integrator data bindings are configured as follows.

Read more about it in the Tom Sawyer Perspectives documentation.

 

Element Bindings

Element bindings correspond to:

  • A URL, either full or relative to the base URL. The value of a URL is defined by an expression.

  • A request method, one of GET, POST, HEAD, OPTIONS, PUT, DELETE, or TRACE.

  • An optional request body. The value of a request body is defined by an expression.

  • An absolute XPath location path. An XPath location path uses XML tags to identify locations within the XML response.

Attribute Bindings

Attribute bindings correspond to a relative path starting from the location path for element bindings.

See how easy it is to connect to a data source and create visualizations in minutes with Tom Sawyer Perspectives.

Writing back changes to the data source

Tom Sawyer Perspectives supports the full data journey, from data integration, graph visualization, and graph editing to writing changes back to the data source.

Support for most data sources

Tom Sawyer Perspectives can write data, or commit changes, back to a Neo4j, Amazon Neptune, Kuzu, Apache TinkerPop, JanusGraph, or OrientDB database, or to an RDF, Excel, SQL, text, or XML data source.

Automated integration handling

Tom Sawyer Perspectives handles the platform-specific details involved in writing changes back to supported data sources.

Maintains data integrity

During commit, Tom Sawyer Perspectives maintains data integrity.

Conflict detection and resolution

Tom Sawyer Perspectives detects and resolves conflicts during update and commit operations. A conflict occurs when the same object with the same identifier has mismatched values in the data source and in the model.


 

Commit configuration

 

Commit configuration

With Tom Sawyer Perspectives, commit configuration gives you control over which changes are committed. By default, automatic bindings for Neo4j, Neptune Gremlin, Neptune openCypher, Kuzu, OrientDB, and TinkerPop integrators handle commit operations automatically. You can also exclude attributes from being committed, and the system supports regular expressions.

Learn more about the graph editing capabilities in Tom Sawyer Perspectives.

The benefits of data commit are clear

Tom Sawyer Perspectives provides the advantages of model persistence while reducing the complexity associated with writing changes back to source systems.

Real-Time Updates

Real-Time Updates

By writing changes directly to the original data source, you can provide real-time updates without delay. This is particularly important in applications where up-to-date information is critical, such as financial systems, real-time monitoring, or collaborative editing tools.

Data Consistency

Data Consistency

Writing data back helps maintain data consistency. When you update the data source immediately after modifying it, you ensure that all users or components working with that data see the same changes. This prevents data discrepancies or conflicts.

Historical Tracking

Historical Tracking

Committing changes in the original data source can provide a complete historical record of all modifications made to the data. This audit trail can be invaluable for debugging, compliance, and accountability purposes.

Scalability

Scalability

In distributed systems, writing changes back to the original source can distribute the data updates across multiple nodes, improving the scalability and load balancing of your application.

Simplified Recovery

Simplified Recovery

If a commit process fails midway, you can use the recorded changes to recover and reapply updates, ensuring data integrity.

Ease of Collaboration

Ease of Collaboration

When multiple users or systems interact with the same data source, writing changes back simplifies collaboration. Everyone sees the same data state, reducing confusion and conflicts.

Reduced Latency

Reduced Latency

In cases where data retrieval is time-consuming (e.g., fetching data from external APIs or databases), persisting changes locally can reduce latency by avoiding repetitive data fetches.

Easier Rollbacks

Easier Rollbacks

If a change leads to unexpected results or errors, reverting to the previous data state is more straightforward when changes are already stored in the original source.

Adherence to Data Policies

Adherence to Data Policies

In scenarios with strict data governance or compliance requirements, writing changes back to the original data source ensures that data policies and access controls are consistently enforced.

Applying visualization and analysis to your data

Once your integrators are configured in Tom Sawyer Perspectives, the data is loaded into an in-memory native graph model, making it effective and efficient to work with.

Tom Sawyer Perspectives makes it easy to design and configure an end-user web or desktop application that uses the graph model, so users can access data in real time. You can configure easy-to-understand views of the data, including graph drawings, tables, charts, timelines, maps, and more. You can also incorporate powerful analytics into the application so users gain deeper insight into the data.

Watch this video to see how easy it is to configure a dashboard-style layout of views for your Tom Sawyer Perspectives application.

An example application created with Tom Sawyer Perspectives showing a dashboard layout with drawing, map, tree, and chart views.

An example application created with Tom Sawyer Perspectives showing a dashboard layout with drawing, map, tree, and chart views.

Query triple stores and labeled property graphs alike without the need to know SPARQL, Gremlin, or Cypher.

Query triple stores and labeled property graphs alike without the need to know SPARQL, Gremlin, or Cypher.

Eliminate the need to know complex query languages

You may not know every query that users will want to perform. That is why we created the Pattern Matching Query Builder, which simplifies advanced graph pattern searches without requiring users to know SPARQL, Gremlin, or Cypher. The Pattern Matching Query Builder allows users to search for matching patterns in triple stores and labeled property graphs alike.

However, users who know Cypher or Gremlin can enter their own queries directly. Adding this capability to your Tom Sawyer Perspectives application is seamless for developers, making data exploration more flexible and developer integration easier.

Improve data navigation and analysis with load neighbors

Load neighbors is a feature in Tom Sawyer Perspectives that improves the data navigation and analysis experience for end users. It enables users to explore their data more effectively, saving time and allowing them to focus on their most important tasks.

With load neighbors, users can load data incrementally based on their use case by searching for graph patterns through an intuitive graph visualization. As a result, they can gain insights and make faster decisions.

first640

Load data incrementally and gain insights faster with the Tom Sawyer Perspectives load neighbors feature.

Get started today

Contact us for a demonstration and to learn how to integrate and visualize your data with Tom Sawyer Perspectives.

FAQs about data integration

What is data integration software?

Data integration software connects data from multiple sources and makes it available in a unified environment where it can be accessed, managed, and used consistently. Depending on business requirements, these solutions may support ETL, data replication, virtualization, real-time integration, or a combination of approaches to move and synchronize data across enterprise systems.

While many platforms focus primarily on moving data between systems, others also support visualization, analysis, and application development. Tom Sawyer Perspectives combines enterprise data integration with a unified graph model, enabling organizations to visualize connected data, analyze relationships, and build graph-based applications without replacing their existing data sources.

What is the difference between data integration and data federation?

Data integration combines data from multiple sources into a unified model that supports analysis, visualization, and application development. Data federation, on the other hand, leaves the data in its original systems and provides a unified way to access it without physically consolidating the data.

The right approach depends on your architecture, performance requirements, and how the integrated data will be used. In many enterprise environments, organizations use both approaches together to meet different business and technical requirements.

What data sources can Tom Sawyer Perspectives integrate?

Tom Sawyer Perspectives integrates data from a wide range of enterprise sources, including relational databases, graph databases, RESTful APIs, RDF repositories, JSON, XML, Microsoft Excel, structured text files, and SQL databases. It also supports graph platforms such as Neo4j, Amazon Neptune, Kuzu, and Apache TinkerPop-compatible databases.

This flexibility allows organizations to work with existing enterprise data without requiring major changes to their current infrastructure.

Does Tom Sawyer Perspectives support graph databases?

Yes. Tom Sawyer Perspectives supports leading graph databases, including Neo4j, Amazon Neptune, Kuzu, and Apache TinkerPop-compatible databases such as JanusGraph and TinkerGraph.

For supported graph databases, Tom Sawyer Perspectives provides automatic schema generation, query-based data binding, graph visualization, and write-back capabilities, making it easier to build interactive graph applications on top of existing graph data.

Can Tom Sawyer Perspectives write changes back to source systems?

Yes. Tom Sawyer Perspectives supports bidirectional data integration for many supported data sources, allowing changes made within an application to be written back to the original data source.

Write-back capabilities help maintain data integrity, simplify collaboration, support conflict detection, and keep enterprise systems synchronized while users work with connected data through graph visualizations and analysis.

Why use a graph model for data integration?

Data integration brings data from different systems together, but a graph model also captures the relationships between that data. By representing both entities and their connections, a graph model makes it easier to understand complex networks, identify dependencies, and analyze connected information.

By bringing integrated enterprise data into a unified graph model, Tom Sawyer Perspectives enables users to visualize relationships, perform advanced graph analysis, and build applications that work naturally with highly connected data.

How can I determine whether Tom Sawyer Perspectives is the right data integration solution?

The right data integration solution depends on your data sources, scalability requirements, integration architecture, and how you plan to use connected data after integration.

Tom Sawyer Perspectives is designed for organizations that need more than traditional data movement. It combines enterprise data integration with graph visualization, graph analysis, and low-code application development, making it well-suited for applications that depend on understanding relationships across complex connected data.